Platform Why Features Security Score AI Engine AI Coding KYP Hub Pricing Company About Buckler News Contact Français Book Demo →
Part Two · The Gap Between Promise and Delivery

The Enterprise AI Challenge

Why pilots start strong and stall - from the Avianca case to the gap between LLM promise and enterprise delivery.

In 2023, a routine lawsuit against Avianca Airlines became a cautionary tale for the AI era. An attorney filed a legal brief drafted with ChatGPT that cited six fabricated court cases, complete with convincing but nonexistent details. When opposing counsel exposed the errors, the case unravelled - leading to a dismissal, a secondary lawsuit, and global headlines.

Mata v. Avianca Inc. made the risk concrete: AI's tendency to hallucinate false information is not a quirk. It is a critical failure mode that can derail any enterprise relying on unverified AI output. The case is one high-profile example of a systemic problem - for every headline incident, countless organizations are quietly experiencing similar disappointments on a smaller scale.

Over the last few years, businesses have poured significant investment into general-purpose LLMs, hoping to streamline operations and unlock new insight. Real-world returns have fallen short of the hype. Pilots that started strongly fizzled out, yielding inconsistent or limited business value. Concerns are growing about whether these resource-intensive models can deliver a reliable return on investment.

The Core Question

LLMs have tremendous potential. The enterprise challenge is not whether the technology works - it is how to unlock that potential at a reasonable cost, with the reliability and governance that business-critical applications demand.

The root issue runs deeper than any single incident. It is a fundamental misalignment between what general-purpose LLMs promise and what they deliver in enterprise environments. To understand why, we examine three primary failure domains: data quality, technical implementation, and business model. Each undermines LLM effectiveness in organizations in a different way, and together they explain why the implementation gap persists.

Part Three · Output That Looks Right But Isn't

Data Quality Issues

Hallucinations, contradictory outputs, and unverified sources - the failure modes that force manual oversight.

1
Hallucinations
Fabricated or inaccurate output presented with unwarranted confidence

LLMs routinely generate information that is fabricated or inaccurate - a phenomenon known as hallucination. The New York Times reported that "the latest OpenAI systems hallucinate at a higher rate than the company's previous system, according to the company's own tests. The company found that o3 - its most powerful system - hallucinated 33 percent of the time when running its PersonQA test, which involves answering questions about public figures." [1] In a business context, an LLM may confidently produce false financial figures or nonexistent product details, eroding user trust on first contact.

2
Contradictory Outputs
Self-inconsistent answers across and within responses

Even when not hallucinating, LLMs can contradict themselves. Studies show ChatGPT-class models exhibit self-contradiction in 17.7% of open-domain text generations [2] - statements that logically conflict with each other within the same response. This stems from vast and sometimes conflicting training data: a user may receive different answers to the same query, or see the model make a claim that does not align with an earlier output. In a business context, an AI assistant might first advise one compliance policy and later suggest the opposite. The inconsistency undermines confidence wherever the system is deployed.

3
Questionable Sources
No built-in guarantee that underlying training data is authoritative

General-purpose LLMs learn from internet-scale data that can be incomplete, low-quality, or biased. They carry no built-in guarantee that a source is authoritative. An LLM may surface outdated or incorrect information from its training corpus; if the underlying data contains misinformation or contradictions, the model reflects them in its answers. Enterprises risk basing decisions on content that has not been vetted - a stark contrast to conventional business intelligence systems that rely on verified data.

The Practical Effect

Current LLMs cannot be trusted for high-stakes enterprise applications without extensive checks. Hallucinations and inconsistencies require manual oversight or secondary validation, which erode the efficiency gains organizations hoped to achieve. Deploying a general LLM "as-is" means accepting an uncomfortably high level of risk.

Parts Four & Five · Implementation and Business Model Barriers

Technical & Business Model Barriers

Inconsistent output formats, maintenance burden, incomplete tooling, unpredictable costs, and vendor lock-in.

1
Inconsistent Output Formats
Free-form text generation resists structured, automatable pipelines

LLMs generate free-form text, which can vary each time - a nightmare for systems that expect structured output. An LLM might return a full-sentence answer on one query, a bullet list on the next, and a pre-determined format on a third, even when the task is identical. Our teams have observed that prompt engineering alone typically achieves only ~36% reliability in producing correctly formatted output, forcing developers to write extensive post-processing code or layer on schema-enforcement features. Minor format drift can break automated pipelines, causing constant rework downstream.

2
Maintenance & Tuning Burden
Model drift demands continuous re-tuning to keep outputs stable

Keeping a general LLM deployment working is a continuous burden. Models may perform well on day one, but as corporate data, user behaviour, or external knowledge changes, responses drift. Internal assistants lose accuracy as new software tools are introduced. Prompt configurations that worked initially need to be revised as outputs evolve. Model providers frequently update their APIs or models, which can alter behaviour or require re-integration. Enterprises must dedicate ongoing resources to monitor output quality, update prompts or fine-tunes, and incorporate new data. Treating an LLM as "set and forget" is a common pitfall.

3
Incomplete Tooling
LLMOps tooling for enterprise integration is still maturing

The surrounding ecosystem for LLM deployment (LLMOps) is still maturing. Integrating an LLM with existing enterprise systems - ERP, CRM, databases - rarely has a plug-and-play solution, and unexpected issues around API payload limits, input sanitization, and security requirements demand custom integration code and infrastructure. Robust tools for versioning prompts, monitoring model decisions, and ensuring compliance are only beginning to emerge. Many organizations end up cobbling together their own frameworks for logging, auditing, and fail-safes because out-of-the-box support is limited. This "assembly required" nature translates to higher implementation cost and complexity for IT.

Net Effect

Deploying a general-purpose LLM in an enterprise setting comes with significant engineering overhead. Inconsistent outputs and constant tuning erode the efficiency gains, while tooling gaps make it difficult to incorporate AI into existing workflows seamlessly. Projects routinely exceed their initial cost plans - and feed directly into the next failure domain: the business model.

4
Unpredictable Costs
Usage-based pricing makes ROI difficult to plan and control

The expense of running LLMs is volatile and hard to control. Most providers charge on usage - token or API-call-based pricing - which means costs scale directly with how heavily employees or applications use the model. Enterprises have repeatedly encountered situations where an AI feature becomes popular and token usage spikes far beyond budget. Self-hosting is no refuge: large models demand powerful and expensive hardware. As day-to-day workflows integrate AI, the aggregate cost "per query" accumulates quickly, sometimes with diminishing returns. Budgeting for an LLM project is tricky - estimates are possible, but actual needs may exceed predictions, and pricing schemes may change. Cost unpredictability makes it difficult to plan ROI and can turn an AI initiative into an unplanned financial drain.

5
Vendor Lock-In and Stability
Dependence on a single external AI vendor introduces strategic risk

Relying on an external AI vendor's model - OpenAI, Google, a startup, or otherwise - introduces strategic risk. If the chosen vendor faces an outage, a policy change, or exits the market, the enterprise's AI capabilities can be disrupted overnight. There is also lock-in risk: switching to another model may require significant rework, and executives worry about the stability of AI vendors in a fast-evolving market where today's leader can become tomorrow's laggard. Trusting a third party with proprietary data through API calls also raises compliance and security questions. Long-term risks such as vendor instability or lock-in have become part of the calculus for every LLM adoption decision. No CIO wants to discover that a mission-critical system breaks because an API was deprecated with little notice.

The Pattern

These business model issues highlight why so many enterprises remain hesitant to fully embrace general-purpose LLMs. Uncertain cost structure and external dependencies conflict with the predictability and control that enterprise software typically demands. For C-level stakeholders, an AI solution must be not only innovative, but also financially and operationally predictable.

Part Six · Architecture

An Enterprise-Grade Architecture

Four purpose-built components that replace probabilistic output with deterministic, enterprise-grade behaviour.

An enterprise-grade platform is engineered specifically for enterprise needs. Instead of relying on a monolithic black-box model, it combines specialized components that work in concert to deliver reliable, actionable intelligence. The architecture centres on four components, each with a distinct role, plus the deployment model that ties them together.

1. Pattern Discovery Engine
A pattern-mining module that ingests and analyzes the organization's own data - documents, databases, logs - to discover meaningful patterns and relationships. The engine acts as a curated knowledge base so the platform operates on verified, high-quality information rather than the open internet. Because every insight is grounded in data the business already trusts, hallucinations are dramatically reduced, and continuous updates keep the knowledge current.
2. Insight Generation Framework
Sits on top of the Pattern Discovery Engine and constructs insights in a consistent, usable format. Where a general-purpose LLM might return a verbose paragraph or an unpredictable structure, the framework applies templates and business rules to produce deterministic outputs - a pros/cons list, a summary report, a JSON snippet ready for an API. Output format is standardized, so integration with dashboards and downstream software is seamless.
3. Real-Time Pattern Recognition
Continuously monitors incoming data - live sales data, market feeds, user queries - and recognizes emerging patterns or anomalies as they happen. The platform updates knowledge and adjusts output on the fly. This lowers the need for manual model re-tuning and improves stability: the platform is less likely to produce outdated advice because it has not seen new data, addressing the model drift issue that plagues static LLM deployments.
4. Business Intelligence Translation
A built-in translation layer between raw AI output and business-level intelligence. Integrates directly with existing BI tools, dashboards, and workflows, so insights are actionable by default. Handles compliance and governance tagging, so every insight carries traceability - source data, confidence level - which is critical for enterprise settings.
5. Deployment Model

The platform deploys in the enterprise's own cloud or on-premises, giving full control over data and cost. Together, the four components deliver advanced AI without the hallucinations, erratic behaviour, hidden costs, or vendor lock-in that characterize general-purpose solutions.

Parts Seven & Eight · Comparison and Conclusion

The Path Forward

Head-to-head comparison, and how to move from pilot to production without inheriting LLM failure modes.

The table below summarizes how a purpose-built enterprise architecture addresses each failure domain covered in this paper, in contrast to typical general-purpose LLMs. The enterprise-grade improvements are concentrated in reliability, maintainability, and cost predictability.

Domain General-Purpose LLMs Purpose-Built Architecture
Data Quality Hallucinations (15-20% of answers incorrect at enterprise scale); self-contradicting outputs; unvetted internet-scale sources Factual, pattern-verified answers; consistent outputs (no self-conflict); uses high-quality enterprise data only
Technical Unpredictable output formats; requires constant prompt tuning; ongoing maintenance & drift issues Structured, deterministic outputs; minimal upkeep with real-time learning; full testing and integration coverage
Business Uncertain usage-based costs; dependence on external vendor; data, security & compliance risks Predictable, fixed cost model; dedicated enterprise support; secure, in-house deployment

Each row in the right-hand column maps to a specific component: data-quality gains come from the Pattern Discovery Engine; technical gains from the Insight Generation Framework and Real-Time Pattern Recognition; business-model gains from the deployment model surrounding the Business Intelligence Translation layer. Each gain is architecturally defensible rather than a prompt-engineering workaround.

The limitations of general-purpose LLMs in enterprise contexts are not superficial. They are structural, and they compound as deployments scale. Closing the gap requires a different architecture, not better prompts. A purpose-built platform represents a fundamental shift in approach: from probabilistic language models to a purpose-built enterprise architecture designed specifically to address the data quality, technical implementation, and business model challenges that have hindered LLM adoption.

By integrating the Pattern Discovery Engine, Insight Generation Framework, Real-Time Pattern Recognition, and Business Intelligence Translation components, a purpose-built architecture delivers the transformative capabilities of advanced AI - without the hallucinations, integration complexity, or unpredictable costs that plague general-purpose solutions. For executives and technology leaders seeking sustainable value from AI investment, this approach offers a path forward that aligns with enterprise requirements for accuracy, reliability, and measurable ROI.

A note on scope: this page reproduces the findings of the research paper Quantifying the Enterprise AI Gap, for readability on the web.
References
  1. The New York Times, OpenAI's New Reasoning AI Models Hallucinate More. Source document
  2. Mündler et al., Self-Contradictory Hallucinations of Large Language Models. Source document