The Case for Building Your Own Small AI Tools
Most organizations have a list of workflows nobody has gotten around to fixing. Not because the fix is technically complex, but because the cost of building anything custom used to be prohibitive. You needed a developer, a project spec, a QA process, and weeks of calendar. For a tool that would save three minutes per task, the math never worked.
That math has changed. The question is whether your decision-making framework has changed with it.
This post is about building small, narrow AI tools internally: when it's worth it, when it isn't, and how to make the call without wasting a quarter on the wrong answer. It's also about the organizational conditions that determine whether anything you build actually gets used.
The Failure Mode Nobody Talks About
By far the most common AI investment mistake right now isn't buying the wrong tool. It's buying a broad tool and expecting it to fix a vague problem.
A MIT Sloan summary of a BCG study run in partnership with Harvard and Wharton found that consultants using GPT improved performance by 38%, and by 42.5% when they were given a task overview alongside the AI access. The same study found that AI used outside its capability boundary reduced performance by 19 percentage points. That's not a rounding error. That's a net loss.
The implication is clear: AI works when it's matched to a specific, bounded task with clear inputs and outputs. It fails when it's deployed against ambiguous work with fuzzy success criteria.
Building small internal tools rather than buying large platforms and hoping they fit addresses this directly. A small tool built for one specific workflow forces you to define that workflow. The act of building is also the act of scoping. And scoping is where most AI deployments fall apart.
What Has Actually Changed in the Build vs. Buy Decision
For most of the last decade, "build vs. buy" was a straightforward calculation for small and mid-sized organizations: buy, always. Building required developer time, infrastructure decisions, and ongoing maintenance that only made sense at scale.
Two things have changed that calculation.
First, the cost of making a narrow tool is a fraction of what it used to be. AI-assisted coding tools, sometimes called vibe coding, have lowered the technical barrier to the point where non-developers can ship working internal tools. Patrick Gilbert covers this in Never Always, Never Never, describing how a team at AdVenture Media built functional web applications during a hackathon run alongside full client workloads. The tools they built were narrow, task-specific, and functional. That's the bar.
Second, off-the-shelf tools have proliferated to the point where evaluating them has become its own cost. Most are mediocre. A small fraction are genuinely useful. Knowing the difference takes time, and even the good ones rarely do exactly what you need for your specific workflow, with your specific data, inside your specific permissions structure.
For workflows that are unique to your organization, narrow, and data-accessible, building is often faster and cheaper than buying and adapting.
When to Build, When to Buy
Research and practical evidence point toward the same operating principle: build when the scope is tight and ownership is clear. Buy when the problem is broad or the workflow changes often.
Here's how to think through the decision:
Build when:
- The workflow is specific to your organization and unlikely to be served well by a generic tool
- The data the tool needs is already internal and reasonably clean
- The tool's function is narrow: drafting, summarizing, classifying, routing, or retrieving from a known document set
- One person or team can own it
- If it fails or produces a bad output, the consequences are reversible
Buy when:
- The problem is broad or touches multiple departments
- The workflow changes frequently
- The data is fragmented, heavily permissioned, or lives across systems
- Maintenance ownership is unclear
- The output directly affects customers, payments, compliance, or other high-stakes decisions without human review
A clear practical boundary: if a tool only reads approved internal content and produces a recommendation that a human acts on, it's often a viable prototype. If it writes back to core systems, moves money, changes customer records, or makes access decisions, it is not a one-day prototype and should not be shipped without a formal review process.
Risk boundaries, not philosophy, drive that distinction. The AI Double Helix framework distinguishes between Internal Efficiency gains and External Value creation, and the risk profile for each is different. Internal tools that produce recommendations are generally safer territory than tools that take autonomous action with external consequences.
What's Actually Buildable Now
To make the build decision concrete, it helps to be specific about what current AI tools can and cannot do reliably.
Today's tools are good enough for internal assistants that:
- Search a limited, curated document set and return relevant passages
- Draft summaries, first-pass emails, and initial analyses from structured inputs
- Classify, route, and extract information from semi-structured text (support tickets, form responses, feedback)
- Help employees navigate standard operating procedures or internal policy documents
- Resize images, reformat data, or perform other repeatable transformation tasks
Current models are not reliable enough to autonomously manage end-to-end processes that involve ambiguous judgment, hidden dependencies, or irreversible actions. Strengths lie in retrieval, drafting, triage, and recommendation, with a human making the final call on anything sensitive.
A worked example worth considering from Never Always, Never Never: an internal "policy copilot" that answers employee questions from a curated HR or operations handbook, cites the relevant internal document, and drafts a support ticket for human review when it's uncertain. Narrow scope. Read-only data. Easy to measure. Reversible if wrong. That's a tool worth building.
A counterexample: an AI assistant that automatically approves refunds or changes customer billing records. That tool touches money, requires durable auditability, and creates high-cost failure modes if it's wrong. Don't build that in a sprint.
The Organizational Conditions That Determine Whether It Works
Ethan Mollick at Wharton has documented that AI benefits are not evenly distributed across tasks or workers. In a consulting-task study he summarized, consultants using AI completed 12.2% more tasks, were 25.1% faster, and produced results rated more than 40% higher in quality. He also documented that lower-performing workers tended to benefit more from AI assistance than higher performers.
What that finding implies for internal tools: the tool is rarely the bottleneck. Organizational conditions are.
BCG's study found that performance improved more when workers received both AI access and a clear task overview than when they received AI access alone. That's a useful organizational rule. Training should be "here's how to use this tool to do this specific job," not "here's a general AI literacy session." Specificity in training produces results. Generality produces compliance theater.
Across the evidence, one pattern stands out: AI ownership should sit close to the business process, not exclusively in central IT. The team that uses the tool should own it, define its success criteria, and have accountability for maintaining it. Where that condition isn't met, tools stall in the pilot-to-production gap, which is a recurring failure mode with documented roots in unclear ownership, weak process integration, and poor data access.
If you want to understand where your organization sits on this spectrum before you build anything, the AI Maturity Ladder is a useful diagnostic. Dabblers and early Practitioners tend to struggle with internal tools not because the tools are hard to build, but because ownership and process integration haven't been defined yet.
The Compound Learning Problem
A timing trap derails most internal AI initiatives. Your first tool will take longer than tools you build later, because you're learning the tools themselves while you're building. That learning compounds. The second tool takes less time. The tenth tool takes a fraction of the first.
Don't wait until you have the perfect use case to start building. Start with something real but low-stakes. A tool that only you use, or that only affects internal processes, or that handles a workflow where being wrong has limited consequences. A Procter & Gamble field experiment involving 776 professionals found that AI-enabled groups worked 12 to 16% faster and produced longer, more detailed outputs. But those gains show up at scale, once people are fluent with the tools. Early investment is in fluency, not immediate output.
Running an internal hackathon or a bounded sprint before assigning AI tools to critical workflows addresses this directly. You're buying learning, not just output. And the learning is what makes every subsequent tool cheaper and faster to build. We wrote about how to structure that kind of event in how to run an AI hackathon that isn't theater.
Compound learning is also why organizational posture matters. If people only have permission to experiment in the margins of an already-full workload, learning stays fragmented. Structured time for building, even brief, changes the trajectory.
Build vs. Buy Decision Audit
Use this before committing resources to either path. It's designed to be run by the person closest to the workflow, not by IT or leadership in the abstract.
---
Internal AI Tool Decision Audit
Copy and fill in for any candidate workflow:
1. Define the workflow
What is the exact task this tool would perform? (Be specific: "Draft a response to a customer support ticket based on our FAQ document" is a workflow. "Help with customer service" is not.)
Workflow: _______________
2. Scope check
- [ ] The tool's task has a clear start and end point
- [ ] The output is a draft, recommendation, or summary (not an autonomous action)
- [ ] A human reviews the output before it affects anything external
- [ ] The tool reads from a defined, accessible data source
If any box is unchecked, the tool is probably out of scope for a small internal build. Revisit the definition or move to the buy column.
3. Data check
- Where does the data the tool needs live? _______________
- Is that data accessible without complex permissions? (Yes / No / Partially)
- Is that data reasonably clean and consistent? (Yes / No / Partially)
If No or Partially on either: add a data preparation phase before building, or reconsider.
4. Ownership check
- Who is the single owner of this tool? _______________
- Does that person have time to maintain it? (Yes / No)
- Is the workflow stable enough that the tool won't need constant rebuilding? (Yes / No)
If No on either: do not build until ownership is resolved.
5. Failure mode check
- What happens if the tool produces a wrong output?
- Is that consequence reversible? (Yes / No)
- Does the output touch payments, customer records, compliance, or irreversible actions? (Yes / No)
If the consequence is irreversible or touches high-risk domains: do not ship without a formal review process, regardless of how well the prototype performs.
6. Build vs. buy call
- Does a good off-the-shelf tool already exist for this exact workflow? (Yes / No / Unknown)
- If yes: what's the estimated time to evaluate and configure it versus build it? Evaluation: ___ hours / Build: ___ hours
- Is the workflow unique enough to your organization that a generic tool is unlikely to fit without major adaptation? (Yes / No)
Decision heuristic:
- Narrow scope + internal data + one owner + reversible failure = strong build candidate
- Broad scope + fragmented data + unclear ownership + irreversible failure = buy or pause
7. Training plan
What specific instruction will users receive? (Not "AI literacy," but "here's how to use this tool to do this specific job")
Training format: _______________
Who delivers it: _______________
Success metric: _______________
---
This last section matters more than most leaders expect. BCG's finding that task overview paired with AI access outperformed AI access alone is a direct argument for writing down exactly how the tool should be used before you ship it. The tool doesn't train itself. Someone has to do the last mile.
The One Thing to Do First
Pick one workflow in your organization that produces a draft, summary, or recommendation for internal use, where the data already exists in a document or spreadsheet, and where a wrong output has limited consequences. Assign one owner. Build a prototype this week, not next quarter. Run it for two weeks, measure whether it saves time or improves output quality, and then decide whether to expand or stop.
Don't start with your most important workflow. Start with one that's real enough to teach you something, low-stakes enough that failure is fine, and specific enough that you can tell whether it worked.
Compound learning starts there.
Patrick Gilbert is the CEO of AdVenture Media and author of Never Always, Never Never and the bestselling Join or Die. He has been ranked among the top 5 PPC experts worldwide and has delivered keynotes at Google events across three continents.
More about Patrick →Enjoyed this?
Subscribe for more articles on strategy, AI, and what's actually working in marketing.
No spam. Unsubscribe anytime.