AI & Decisions

When the system should refuse to recommend

A recommendation engine that always answers is not safer. Useful systems know when the evidence is too thin and say so.

Effective
Last updated
Reading time
6 min

In June 2023, a Manhattan federal judge sanctioned two lawyers and their firm $5,000 for filing a brief with cases that did not exist. ChatGPT had produced the fake cases. The output looked polished. The citations looked real. The system answered when it should not have.

That is the problem with many AI products and recommendation engines: they are built to always answer.

A system that always recommends is not more useful. It is more dangerous. Sometimes the correct output is "not enough evidence."

Confidence laundering

Bad recommendations often look like good recommendations.

They have the same UI. The same tone. The same formatting. The same confidence. A user cannot tell whether the system is working from strong evidence, weak evidence, stale data, out-of-scope data, or no real evidence at all.

The presentation launders confidence.

That shows up in legal, healthcare, finance, hiring, customer support, and growth work. A system can recommend scaling a campaign when spend data is missing. It can recommend killing a test before enough conversions have arrived. It can recommend a refund policy that does not apply. It can rank a sales account when the last activity is 90 days old.

The user pays for the system's unwillingness to say no.

Why always-on recommendations are especially risky in growth

Growth recommendations touch money quickly.

A recommendation to scale a campaign can move budget today. A recommendation to kill a test can stop a winner before the signal matures. A recommendation to promote a variant can change a landing page, email, offer, or checkout for thousands of visitors. A recommendation to rebalance channels can move spend away from a channel that looks weak only because attribution is broken.

That is why recommendation quality is not only about model accuracy. It is about operating context.

The system needs to know whether spend data is current, whether revenue is joined, whether refunds are inside the window, whether the cohort is large enough, whether a lifecycle rule has fired, whether a spend cap exists, whether the campaign is still learning, and whether the metric is causal or only directional. If those inputs are missing, the recommendation should not be dressed up as insight.

In many teams, the worst recommendation is not an obviously wrong one. It is a plausible one that arrives with the same confidence as a strong one. That is how weak signal becomes budget movement.

The four-part contract

A recommendation worth trusting should carry four things.

Applicability. Who is this recommendation for? Which campaign, customer, audience, market, time window, or policy does it apply to?

Confidence. How strong is the evidence? Use a plain class like high, medium, low, or insufficient. Do not hide weak evidence behind a precise-looking score.

Guardrails. What rules limit the recommendation? Spend caps, lifecycle rules, approval thresholds, margin floors, privacy limits, policy constraints, or minimum sample sizes should be named.

The right to abstain. When the evidence is too thin, the system should return no recommendation and explain why.

The fourth part is the load-bearing one. Without it, the system is forced to produce a weak answer.

A concrete budget example

Imagine a Meta campaign with $18,000 of spend this month and a platform ROAS above target. The obvious recommendation is "increase budget."

But the system should check more than platform ROAS:

  • Is spend ingestion current?
  • Is revenue joined to orders and payments?
  • Are refunds or cancellations inside the read window?
  • Is MER improving or only platform ROAS?
  • Is CAC below the target for the audience?
  • Is the campaign past the learning window?
  • Is frequency rising into fatigue?
  • Is there a spend cap or margin floor?
  • Did a related offer or landing page test change at the same time?

If the answers are strong, "scale" may be right. If revenue is missing, the correct output is "do not recommend because revenue is unavailable." If frequency is high and marginal return is falling, the correct output may be "review saturation before scaling." If the campaign is under target but the audience is strategic, the correct output may be "extend test."

The value is not the action label. The value is the discipline that decides whether any action should be recommended.

What research shows

Abstention is not a vague product preference. It is a measurable technical capability.

Selective classification research shows that systems can answer fewer cases and become much more reliable on the cases they do answer. Work by Geifman and El-Yaniv showed this trade-off clearly. Work by Mozannar and Sontag extended the idea to systems that can decide when to defer to a human expert.

The lesson for product teams is practical: coverage is not always the goal. A system that answers 80% of cases with high reliability can be more valuable than a system that answers 100% of cases and quietly fails on the hardest 20%.

What this means for growth teams

Growth teams do not need an AI that always has an opinion.

They need a system that can say:

  • scale this campaign because spend, revenue, and margin are all verified
  • kill this campaign because it is over spend cap and below target
  • promote this winner because the test crossed the decision rule
  • do not recommend because spend data is missing
  • do not recommend because the cohort is too small
  • do not recommend because the attribution chain is weak

Those last three are not failures. They are useful outputs.

What Lyberty does

Lyberty's recommendations include the campaign, audience, time window, confidence class, guardrails, and evidence behind the call.

The recommendation engine can emit actions like kill campaign, scale campaign, promote winner, decrease budget, investigate underperformance, rebalance channel mix, fix saturation block, or review target gap. But it only recommends when the inputs support the decision. If the upstream signal is missing or a guardrail fails, the system abstains.

That is important because Lyberty's recommendations can affect real budget decisions. A polished guess is not good enough.

What to ask

  1. What percentage of cases does the system refuse to answer?
  2. What reasons can it give for refusing?
  3. Can those reasons be inspected by an operator?
  4. Are recommendations tied to spend caps, targets, lifecycle rules, and evidence?
  5. Does the product treat abstention as a first-class output or as an error?

If the system has no abstention path, it has no real confidence model.

Sources

  1. Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. June 22, 2023). https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1:2022cv01461/575368/54/
  2. Dahl, M., Magesh, V., Suzgun, M., and Ho, D. E. (2024). "Large Legal Fictions." https://hai.stanford.edu/news/hallucinating-law-legal-mistakes-large-language-models-are-pervasive
  3. Moffatt v. Air Canada, 2024 BCCRT 149. https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
  4. McKinsey & Company (2024). "The state of AI in early 2024." https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-2024
  5. Gartner (2024). "Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025." https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025
  6. Geifman, Y. and El-Yaniv, R. (2017). "Selective Classification for Deep Neural Networks." https://arxiv.org/abs/1705.08500
  7. Mozannar, H. and Sontag, D. (2020). "Consistent Estimators for Learning to Defer to an Expert." https://proceedings.mlr.press/v119/mozannar20b.html
  8. Rudin, C. (2019). "Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead." https://www.nature.com/articles/s42256-019-0048-x
  9. Wong, A., Otles, E., Donnelly, J. P., et al. (2021). "External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients." https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307
  10. Vectara (2025). "Hallucination Leaderboard." https://github.com/vectara/hallucination-leaderboard