Choose the action before choosing the technology
“We need an AI agent” is not yet a workflow requirement. “Every suitable enquiry needs an owner and a booking option” is closer. It names a result you can inspect.
Rule-based automation follows explicit conditions. AI-assisted automation can interpret varied language and draft context-aware text. Neither label tells you whether the complete workflow is reliable. Reliability depends on the inputs, allowed actions, stored state, failure handling and review process.
For lead handling, a mixed design is often a reasonable starting point: rules control state changes and permissions, while a model assists with language tasks. That is a design recommendation, not a claim that one architecture wins every comparison.
Divide the workflow into small decisions
List what happens after an enquiry arrives. The list may include recording the message, assigning an owner, identifying the requested service, asking for a missing detail, offering a booking link and stopping follow-up after booking.
Some steps have clear inputs and outputs. Others need interpretation. An explicit service dropdown can route by a rule. A paragraph mentioning several problems may need a person or a model-assisted classification.
The same workflow can contain both. You do not need to make the calendar, consent check and message interpretation all depend on one model response.
| Task | Starting approach | Reason |
|---|---|---|
| Store a form submission | Deterministic workflow | The expected fields and destination are defined |
| Stop follow-up after booking | Rule using confirmed booking state | The stop condition should be explicit |
| Interpret a free-text enquiry | AI suggestion with uncertainty handling | Language may vary beyond a small keyword list |
| Send verified appointment details | Template with current record values | Creativity adds little to time and location |
| Answer unusual pricing requests | Human review, optionally with a draft | The reply may create a commercial commitment |
| Summarise a conversation for an owner | AI-assisted summary checked against history | Useful compression still needs factual accuracy |
Keep decisions and actions separate
Suppose a model suggests that a lead is suitable. The next action should still check that the required information exists and that booking is permitted. A persuasive sentence should not bypass a missing record or a request for a person.
Define a small output contract: requested service, missing information, proposed route and supporting text from the enquiry. If the output is incomplete or contradicts the known record, send it to review rather than guessing.
Avoid allowing the model to invent product details. Provide a maintained source for prices, supported services and exclusions. When the answer is missing, the correct next step is clarification or handoff.
Test with difficult inputs, not only the demo
Build a small, anonymised or synthetic test set containing clear fit, clear non-fit, unclear wording, multiple requests, contradictory answers, a booking already made and a request to stop.
Write the expected action before running the workflow. Otherwise a fluent answer can persuade you that an incorrect route is acceptable. Inspect the stored result and actual permitted action, not only the text on screen.
Repeat tests after changing the model, prompt, source material or routing rules. Compare error types: wrong classification, invented information, lost context, repeated questions and actions taken after a stop condition.
Do not claim accuracy from a handful of examples. A small set helps expose obvious failures. It does not estimate every failure rate in production.
Latent Space's interview with AIUC discusses the gap between a successful demo and testing adversarial failures. For a lead-handling pilot, turn that question into a practical test: what happens when a customer contradicts the record, asks for an unsupported promise or requests that the automation stop? This is a proposed test approach, not evidence that Rainlight has passed an independent certification.
Include operating cost and human work
A cheap model can be expensive if staff must repair its mistakes. A rigid form can be expensive if it pushes useful enquiries into an abandoned path. Record software costs, review time, exception volume and the consequence of an incorrect action.
Choose thresholds based on the business's own tolerance and task. There is no universal model confidence score that makes a commercial answer safe to send. Many systems report a score without demonstrating that it predicts correctness.
Start with drafts or suggestions if the team's rules are unsettled. Move a narrow action to automatic execution only after its inputs, boundaries and verification are clear.
Define what improvement would look like
The target is a better handled enquiry, not a more impressive agent demo. Measure useful-response delay, correct routing, unresolved exceptions, suitable bookings and staff correction time.
If a rule solves the problem clearly, use it. If language variation blocks the rule, test an AI-assisted step. If neither can resolve the decision reliably, give a person the context and ownership. Rainlight's handoff guide helps make that last route an operating process rather than a vague fallback.