Never Always, Never NeverNever Always,
Never Never
The BookAI CreativeAI ProjectsChatNewsletter
Buy the Book
Never Always,
Never Never.

Strategic Marketing in an AI World.
By Patrick Gilbert.

Explore

  • The Book
  • The Author
  • Free Chapter
  • Buy
  • Resources

Connect

  • Newsletter
  • Learn
  • AI Projects
  • Blog
  • Chat with the Book
  • Subscribe

Stay Connected

Get updates, bonus frameworks, and new AI project showcases.

© 2026 Patrick Gilbert. All rights reserved.

AdVenture MediaContact
AI Enablement9 min readSeptember 27, 2026

Context Engineering: Your AI Output Is Only as Good as What You Feed It

Patrick Gilbert

Patrick Gilbert

CEO of AdVenture Media. Author of Never Always, Never Never.

Half of large enterprises with AI in production cannot prove whether it works. That is the finding from the Plug and Play 2026 Enterprise AI Strategy Pulse Survey, and it is probably the most damning sentence in any AI adoption report published this year.

These are not companies that haven't tried. The same survey found 74% of large enterprises have at least one AI solution running in production. They have the tools. They have the spend. They do not have the results. The gap between deployment and value is not a technology problem. It is an information problem.

What you give the model to work with determines everything. If that information layer is missing, stale, fragmented, or badly scoped, no amount of prompt rewording will fix it. Context engineering is built on this argument, and getting it right is the single most important operational change most teams can make.

What Context Engineering Actually Means

Prompt engineering is about phrasing. Context engineering is about what the model can see.

A better prompt might help you get a more coherent sentence. Better context determines whether the model is working from your actual pricing policy, your current product catalog, and your approved compliance language, or whether it is guessing based on patterns from its training data.

The distinction matters because the failure mode for most enterprise AI is not that the model is dumb. It is that the model is working blind. It gives plausible-sounding answers that do not reflect your organization because it has no access to your organization. Anthropic's published documentation on contextual retrieval makes this concrete: their system prepends chunk-specific explanatory context before embedding and indexing documents, so that when the model retrieves information, it retrieves relevant, scoped evidence rather than generic matches. The mechanism is technical, but the principle is simple. The model performs better when it can retrieve current, specific, source material rather than infer from a vague instruction.

In practice, this translates directly for business teams: build the information environment first, then write the prompt.

The Jagged Frontier Problem

Wharton professor Ethan Mollick's summary of the BCG and Harvard field experiment is the most useful piece of AI adoption research published in recent years, and it contains a warning that most organizations have not fully absorbed.

Researchers randomized 758 consultants into three groups: no AI, AI only, and AI plus instruction. Results for tasks inside the model's capability frontier were striking. Consultants using AI completed 12.2% more tasks, finished 25.1% faster, and produced more than 40% higher-quality results. Those numbers get cited constantly.

What gets cited less is what happened outside the frontier. For tasks beyond what the model could reliably handle, AI users were 19 percentage points less likely to produce correct solutions than those working without AI. The tool that dramatically improved performance on the right tasks actively degraded performance on the wrong ones.

Mollick calls this the "jagged frontier." The boundary between what AI handles well and what it handles badly is not visible to most users, and it does not map neatly to how hard a task looks to a human. Two tasks that feel equally complex to the person doing them can sit on opposite sides of that frontier.

For teams, this is a context engineering problem as much as it is a training problem. Part of what the "AI plus instruction" group received was guidance on where the model helps and where it breaks. That is not a prompting tip. That is an operating model. Knowing which tasks to hand off and which to keep requires that the organization has mapped its own work against the frontier, which means understanding how AI works well enough to make that call.

Patrick Gilbert's book Never Always, Never Never frames this through the 4x2 Model of Work. The model collapses the old approach of working solo, copiloting, or delegating into two modes: Copiloting, where the human stays in the driver's seat and AI acts as a capable navigator, and Delegating, where the human hands the work over entirely and shifts into an editorial role. Both modes apply to different types of tasks, and getting the assignment wrong is expensive. Delegating a task that sits outside the frontier produces the BCG experiment's worst result: confident-sounding wrong answers that the human accepts because they trusted the system.

Context engineering reduces that risk by giving the model better raw material and by scoping which tasks the model is authorized to handle.

Why the Knowledge Base Is the Actual Product

Most organizations treat their AI implementation as a software project. Install the tool, train the team on prompting, measure outputs. The context layer, meaning the actual documents, policies, procedures, and source material the model works from, gets treated as a secondary concern.

This is backwards. The knowledge base is the product. The model is the engine.

Consider what happens when a team member asks an AI assistant a question about your return policy, your data handling procedures, or your pricing structure. If those documents are not in the system, the model either says it does not know, which is the honest answer, or it generates a plausible-sounding response based on patterns, which is the dangerous answer. In a production environment where people trust AI outputs, the second outcome causes real problems.

Durable results come from combining three changes: a maintained knowledge layer, clear task boundaries, and training that teaches people the operating model rather than just the tool. The BCG experiment's high-performing group did not just get access to a model. They got guidance documents and prompt-training materials alongside the tool. That package, not the model alone, drove the productivity gains.

This connects to what the AI Double Helix framework describes as Internal Efficiency: the foundational strand of an AI-first organization. Before you can use AI to create external value for customers, you need the internal information architecture that makes the model reliable. Companies that skip this step build on sand. They get impressive demos and inconsistent production results, which is exactly what the Deloitte 2026 survey found: only 25% of respondents had moved 40% or more of their AI pilots into production, even though 54% expected to reach that threshold within three to six months.

That gap between those numbers is not a capability gap. It is a readiness gap. Data that is fragmented, undocumented, or ungoverned cannot be turned into reliable context, and without reliable context, pilots do not survive contact with production.

What Good Context Engineering Looks Like in Practice

Mechanics here are not complicated. What makes this hard is the organizational discipline required to maintain it.

A working context layer for a business team has six components:

Source documents. Approved, current versions of policies, product specifications, pricing, process guides, and legal constraints. The model works from what you give it. If you give it a policy document that is two versions out of date, you get answers based on the old policy.

Named ownership. Every knowledge domain needs a maintainer. Not a team, not a department. A named person responsible for keeping that section current.

Retrieval rules. What sources may the model draw from? What is out of scope? What should happen when sources conflict? These rules need to be explicit, not assumed.

Task boundaries. Which tasks are safe to delegate fully? Which require human review before the output goes anywhere? Which are out of scope entirely? The BCG experiment's lesson applies here directly: the frontier is not obvious, and leaving it undefined invites the worst outcome.

Quality checks. How do you know the system is working? Accuracy rate on verifiable outputs, escalation frequency, time-to-resolution, error categories. Vague satisfaction is not a metric.

Update cadence. Knowledge goes stale. Who refreshes which documents, and on what schedule? This is the maintenance discipline that most organizations skip and almost all regret.

At AdVenture Media, shifting toward this kind of operating structure was part of what Patrick Gilbert describes as moving from AI-Forward to AI-First: from using tools opportunistically to building the institutional layer that makes those tools reliable at scale.

To read more about the human side of that transition, the post on building an AI-first culture covers the organizational dynamics that tend to derail it.

How to Tell When Context Is the Problem

Most teams blame the model when the real issue is the information environment. These are the signals that context is your bottleneck, not capability:

  • The model gives different answers to the same question across sessions, because it has no stable source material to anchor to.
  • It performs well on generic tasks but fails on anything company-specific.
  • Errors disappear when you manually paste the relevant document into the prompt.
  • Performance improves when retrieval is added, but not when the prompt is rewritten.
  • Your team keeps asking for better prompts when the real issue is that the right documents are not in the system.

That last signal is the most common. Prompting tips are easy to search for and easy to share. Building and maintaining a knowledge base requires ownership and process. Organizations default to the easier intervention and wonder why results stay inconsistent.

A deeper look at what happens when AI rollouts stall at the organizational level is available in the post on why AI rollouts fail, which covers the structural patterns that cause pilots to die before they reach production.

Context Engineering Audit: Copy and Use This

Run this audit on any AI use case your team has in production or is piloting. Answer each question. Where you cannot answer, that is your gap.

---

Part 1: Knowledge Layer

Source documents

  • [ ] What documents does the model currently have access to for this use case?
  • [ ] Are those documents the current, approved versions? When were they last updated?
  • [ ] Who owns each document and is responsible for keeping it current?
  • [ ] Are there knowledge gaps where the model is expected to answer questions it has no source material for?

Retrieval rules

  • [ ] Is retrieval set up, or is the model working purely from its training data?
  • [ ] What sources is the model authorized to draw from?
  • [ ] What sources is it explicitly excluded from?
  • [ ] What happens when two sources conflict? Is there a defined tiebreaker?

---

Part 2: Task Boundaries

  • [ ] Which specific tasks is this AI use case handling?
  • [ ] For each task: is it inside the model's reliable capability frontier, or are you assuming it is?
  • [ ] Which outputs go directly to end users or downstream systems without human review?
  • [ ] Which outputs require a human check before they go anywhere?
  • [ ] Is there a documented list of tasks this AI is explicitly not authorized to handle?

---

Part 3: Quality and Measurement

  • [ ] How do you verify that outputs are accurate? Name the specific check.
  • [ ] What is your current error rate on verifiable outputs?
  • [ ] How often does the model escalate to a human, and is that rate what you expected?
  • [ ] What business outcome are you measuring, beyond task completion? (Time saved, error reduction, resolution rate, etc.)
  • [ ] How would you know if performance degraded? Is there a monitoring mechanism?

---

Part 4: Maintenance

  • [ ] Who is responsible for updating the knowledge base when policies or products change?
  • [ ] Is there a defined update cadence, or does it happen reactively?
  • [ ] When did you last review the source documents for accuracy and completeness?
  • [ ] Is there a process for users to flag when AI answers are wrong or outdated?

---

Scoring

Count your unanswered questions.

0-3 gaps: Your context layer is reasonably mature. Focus on measurement and cadence.

4-8 gaps: You have a context problem. Prioritize source documents, ownership, and task boundaries before investing further in prompting or tooling.

9+ gaps: Your AI use case is running on assumptions. Pull it back to a controlled pilot and build the information architecture before scaling.

---

The AI Maturity Ladder described in Never Always, Never Never puts this in organizational terms. Practitioners use AI to solve daily friction points. Architects build the systems that make AI reliable across a team. The audit above is Architect-level work. If your organization is producing Practitioner-level outputs and wondering why performance is inconsistent, this is usually the reason.

Start Here

Pick one AI use case your team is already running. List every question the model is expected to answer. Then ask: what source document contains the correct answer to each question, and is that document in the system?

Every gap you find is a context problem, not a prompting problem. Fix the information environment first. The prompt is the last thing you should be adjusting.

Patrick GilbertPatrick Gilbert

Patrick Gilbert is the CEO of AdVenture Media and author of Never Always, Never Never and the bestselling Join or Die. He has been ranked among the top 5 PPC experts worldwide and has delivered keynotes at Google events across three continents.

More about Patrick →

Enjoyed this?

Subscribe for more articles on strategy, AI, and what's actually working in marketing.

No spam. Unsubscribe anytime.

Keep reading

AI Enablement

Why AI Rollouts Fail: The Real Reasons Behind Enterprise AI Implementation Failure

95% of enterprise AI pilots deliver no measurable ROI. Here's why AI adoption fails, and what leaders can do before the next rollout.

AI Enablement

The AI Maturity Ladder: Where Your Team Actually Sits

93% of data leaders are experimenting with AI. Only 7% have reached enterprise-wide deployment. Here's the framework that explains the gap.

AI Enablement

Copiloting vs Delegating: Which Tasks to Hand Over to AI

Most AI rollouts fail because teams delegate the wrong tasks. Here's the framework for deciding what to hand over and what to keep.