A human hand calibrating one precise red-lined workflow within a larger unactivated technical system.

AI / 12 March 2026

Practical AI starts smaller than most teams expect

The best early AI projects are rarely grand transformations. They are focused workflows where the business can prove value quickly.

Key Takeaways

  1. Select one repeated, bounded workflow with a clear owner and review point.
  2. Establish a baseline and an evaluation set before choosing a model or interface.
  3. Treat governance, security and human accountability as part of the product design.

The first useful AI project inside a business is unlikely to look like transformation. It will look like a recurring task becoming easier to complete, review and improve.

That modest scope is an advantage. It creates a controlled place to learn what the technology can do, where it fails, which data matters and how accountability should work. A broad initiative can hide those questions behind ambition. A bounded workflow forces the team to answer them.

The objective is not to prove that AI is impressive. It is to establish whether a specific system produces enough reliable value to deserve a larger role.

Start with friction, not technology

Do not begin with a model demonstration and search for somewhere to deploy it. Begin with work that already exists.

Good candidates often share five characteristics:

  1. The task happens frequently enough for improvement to matter.
  2. The inputs and expected outputs can be described.
  3. A knowledgeable person already knows how to review the result.
  4. Failure is visible and recoverable.
  5. The workflow has a clear owner.

Examples include finding precedent across an internal knowledge base, preparing a structured first draft from approved source material, classifying inbound requests or extracting fields from a known document type.

Poor first candidates are ambiguous, politically sensitive or difficult to reverse. If success depends on an unspoken expert judgement that the team cannot explain, it will also be difficult to evaluate.

Define the baseline before the prototype

A prototype can feel faster while creating more review work elsewhere. Measure the current workflow first.

Record a simple baseline:

  • how long the task takes from request to accepted output;
  • where people search for source material;
  • which errors cause rework;
  • who reviews and approves the result;
  • what quality looks like in observable terms;
  • which information must never leave an approved boundary.

This turns “AI saved time” into a testable question. It also reveals whether the real constraint is generation at all. Sometimes the better intervention is clearer source content, a better search index, a form or a conventional automation.

Design the smallest complete loop

A useful pilot is small in scope but complete in operation. It includes the source, interaction, output, review and learning loop.

Input

Specify exactly which information the system may use. Remove duplicate, obsolete or contradictory source material before expecting retrieval to solve it.

Transformation

Define what the system is allowed to do: retrieve, summarise, classify, extract or draft. Combining all of these in the first version makes failures harder to locate.

Review

Name the person accountable for accepting, changing or rejecting the output. “Human in the loop” is not a control unless the human has enough context, time and authority to intervene.

Feedback

Capture why an output was rejected or edited. Those reasons become evaluation cases and expose missing context, weak instructions or unsuitable tasks.

This is the minimum viable intelligence loop. Without feedback, the system repeats work. Without review, the business cannot distinguish assistance from ungoverned delegation.

Build an evaluation set

Before tuning prompts, collect a representative set of real examples. Include routine cases, difficult edge cases and inputs that should be refused or escalated.

For each case, define the qualities that matter. Depending on the workflow, those may include:

  • source accuracy;
  • completeness;
  • correct classification;
  • tone and format;
  • appropriate uncertainty;
  • citation of supporting material;
  • escalation when evidence is missing.

Use a scoring method that the domain owner understands. Some criteria can be automated, but expert review is often required for meaning and risk. Keep the set stable enough to compare versions, then add new failure cases as the pilot encounters them.

The point is not to find one universally best model. It is to know whether the chosen configuration is good enough for this workflow, with this data, under these controls.

Governance belongs inside the design

The NIST AI Risk Management Framework organises AI risk work around governing, mapping, measuring and managing. Those activities are useful at pilot scale, not only for enterprise programmes.

Govern

Assign ownership, usage rules and escalation paths. Decide who can change prompts, sources and model settings.

Map

Describe the users, affected people, data flows, intended use and plausible misuse. A system used for internal drafting carries different consequences from one making customer-facing recommendations.

Measure

Test quality, reliability, privacy and security against the evaluation set. Record model and prompt versions so results are reproducible.

Manage

Set thresholds for release, monitoring, rollback and retirement. A pilot should have a stop condition as well as a success condition.

The companion NIST AI RMF Playbook provides voluntary actions teams can adapt to their context. The value is the discipline of making risk decisions explicit.

Treat retrieved content as untrusted input

An assistant connected to documents, websites or messages is exposed to more than factual errors. External content can contain instructions that alter model behaviour.

OWASP identifies prompt injection as a leading risk for LLM applications, including indirect injection through content the system retrieves. Retrieval and fine-tuning do not remove that risk.

Practical controls include limiting tool permissions, separating instructions from retrieved data, validating outputs before actions, constraining destinations, logging tool use and requiring approval for consequential operations. Do not let a successful drafting demo quietly evolve into an autonomous system with broad access.

A six-week pilot shape

A focused pilot can follow this sequence without pretending every organisation moves at the same speed.

Phase Decision
Frame Is the workflow bounded, valuable and owned?
Baseline What happens now, and how will improvement be measured?
Prepare Are source content, permissions and test cases ready?
Prototype Can the complete input-to-review loop work?
Evaluate Does it meet quality and risk thresholds on representative cases?
Decide Scale, revise, contain or stop?

The final phase is essential. A pilot is an evidence-generating exercise, not a commitment to permanent deployment.

Know when to scale

Scale only when the team can answer these questions with evidence:

  • Does it improve the baseline without moving hidden work to reviewers?
  • Are the common failure modes known and detectable?
  • Can users understand the limits of the output?
  • Are permissions and data handling appropriate?
  • Is there an accountable owner for performance after launch?
  • Can the system be monitored, changed and withdrawn safely?

If the answers are weak, a larger rollout multiplies uncertainty. If the answers are strong, the organisation has something more valuable than a demonstration: a repeatable way to identify, evaluate and govern the next workflow.

Practical AI starts small because learning is the first deliverable. Scale should be earned by what that learning proves.

Source Notes

References

  1. AI Risk Management FrameworkNational Institute of Standards and Technology
  2. NIST AI RMF PlaybookNational Institute of Standards and Technology
  3. LLM01:2025 Prompt InjectionOWASP GenAI Security Project

Thinking Into Action

Turn the idea into a useful digital system.

Bring the current challenge, the direction you are considering and the system that needs to exist next.

Start a conversation