In June 2023, a New York lawyer filed a federal brief with six cases that did not exist. The fake cases had names, docket numbers, quotes, and legal reasoning. ChatGPT had produced them. When the lawyer asked the chatbot if the cases were real, it said yes. Judge P. Kevin Castel sanctioned the lawyers and their firm $5,000.
Eight months later, Air Canada's chatbot gave a customer the wrong refund guidance. The airline was held responsible.
Both failures were described as AI hallucinations. That is only part of the story. The deeper failure was that the answer had no reliable receipt. The system could not prove which source supported which claim, what policy allowed the answer, or what record should be reviewed later.
Retrieval-augmented generation, or RAG, solved one problem: find relevant context. Production systems now need to solve the next problem: prove what the answer was based on.
The first RAG wave
The first RAG wave was built around vector search.
The common pattern was simple:
- Load documents.
- Split them into chunks.
- Embed the chunks.
- Store them in a vector database.
- Retrieve likely matches at query time.
- Put those chunks into a prompt.
- Generate an answer.
That pattern was useful. Pinecone, Weaviate, Chroma, Qdrant, Vespa, Elastic, Azure AI Search, LangChain, and LlamaIndex all helped make retrieval easier. Hybrid search, rerankers, contextual retrieval, GraphRAG, Self-RAG, FLARE, and HyDE improved retrieval quality.
But better retrieval is not the same as accountable answering.
The system can retrieve the right documents and still make an unsupported claim. It can cite a passage that does not support the sentence. It can merge two sources incorrectly. It can answer outside the user's permission scope. It can leave no useful audit record.
Why citations are not enough
Most business users do not need academic citations. They need usable proof.
A footnote at the end of an AI answer can create the feeling of proof without actually proving the claim. The answer may cite the right document but the wrong section. It may cite a page that contains related language but does not support the recommendation. It may combine two documents that were never meant to be combined. It may answer from a document the user was allowed to read but apply it to a customer, region, or policy case where it does not apply.
That distinction matters in company work.
If an AI system summarizes a product memo, a loose citation may be fine. If it recommends a refund, a hiring action, a price exception, a campaign budget change, a compliance answer, or a customer-facing promise, a loose citation is not enough. The system needs to show which evidence supports which claim, and which policy allowed the answer to be used.
The real question is not "did the model cite something?" The real question is "could a human reviewer inspect the answer later and understand why it was allowed?"
The three layers
Production RAG needs three layers.
1. Retrieval. Find relevant context. This is where vector search, BM25, hybrid search, rerankers, query rewriting, GraphRAG, and agent tool calls live. This layer is now fairly mature.
2. Evidence. Bind the answer to the source. For each important claim, the system should know which passage, row, field, document, or record supports it. A citation at the end of a paragraph is not enough if nobody can tell which sentence it supports.
3. Policy. Decide what the system is allowed to say. The answer should respect permissions, role, data scope, customer commitments, regulated language, and human-review rules. When the evidence is too weak, the system should refuse to answer or route to review.
Most RAG tools ship layer 1 and leave layers 2 and 3 to the customer.
The operating standard
Reliable AI answers should be built like business records, not like chat messages.
For every important answer, the system should keep three separate things:
The question context. Who asked, what they were allowed to access, what customer, campaign, account, policy, or document set the question referred to, and what time window mattered.
The evidence set. Which passages, rows, files, records, events, or prior decisions were actually used. Not every retrieved document. The specific evidence that supports the important claims.
The decision status. Whether the system answered, refused, routed to review, or produced a draft that still needs human approval.
This is how AI becomes usable inside real workflows. The company can let AI work faster without losing the ability to inspect, challenge, or correct the output. The answer becomes part of the work history instead of another polished paragraph floating in chat.
What a receipt should include
For a business-critical AI answer, the receipt should show:
- the question
- the user and permission scope
- the sources searched
- the evidence actually used
- the specific claims supported by each evidence item
- the policy version in force
- the model and prompt version
- whether the system answered, abstained, or routed to review
That sounds heavy until something goes wrong. Then it is exactly what legal, compliance, support, and the customer will ask for.
Why this matters for company work
If AI is only drafting low-stakes copy, a loose citation may be acceptable.
If AI is helping with refunds, pricing, compliance, hiring, medical operations, financial decisions, customer promises, or budget recommendations, loose citations are not enough. The company needs to know which answer came from which evidence and whether the system was allowed to produce it.
The EU AI Act and NIST AI Risk Management Framework both point toward traceability and human oversight. Buyers will ask for this even before regulators do, because they need to trust the system internally.
What Lyberty does
Lyberty records evidence as part of AI-assisted work.
A retrieval run keeps the corpus scope, evidence references, policy version, model context, and the final output. When the system makes a recommendation or drafts an answer from internal context, it can show what supported the answer. If the evidence is too thin, the system can abstain instead of filling the blank with confident prose.
Lyberty does not claim perfect claim-level proof for every kind of reasoning. When the model goes beyond retrieved evidence, that is marked differently from a grounded answer. The point is to make the confidence visible instead of hiding it behind a polished paragraph.
What to ask any AI vendor
- Show me one answer.
- For each important sentence, show the source that supports it.
- Show which documents the user was allowed to search.
- Show the policy that governed the answer.
- Show the audit record six months later.
If the vendor cannot do that, it built a search-and-answer tool, not a reliable work system.
Sources
- Lewis, P., Perez, E., Piktus, A., et al. (2020). "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS 2020. https://arxiv.org/abs/2005.11401
- Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. June 22, 2023). https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1:2022cv01461/575368/54/
- Moffatt v. Air Canada, 2024 BCCRT 149. https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
- Edge, D., Trinh, H., Cheng, N., et al. (2024). "From Local to Global: A Graph RAG Approach to Query-Focused Summarization." https://arxiv.org/abs/2404.16130
- Anthropic (2024). "Introducing Contextual Retrieval." https://www.anthropic.com/news/contextual-retrieval
- Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H. (2023). "Self-RAG." https://arxiv.org/abs/2310.11511
- Jiang, Z., Xu, F. F., Gao, L., et al. (2023). "Active Retrieval Augmented Generation." https://arxiv.org/abs/2305.06983
- Gao, L., Ma, X., Lin, J., and Callan, J. (2022). "Precise Zero-Shot Dense Retrieval without Relevance Labels." https://arxiv.org/abs/2212.10496
- Andreessen Horowitz (2023). "Emerging Architectures for LLM Applications." https://a16z.com/emerging-architectures-for-llm-applications/
- Pinecone Systems (2023). "Pinecone Raises $100M Series B." https://www.pinecone.io/blog/series-b/
- Cohere (2024). "Cohere Raises US$500 Million Series D." https://cohere.com/blog/cohere-raises-500m-series-d
- MongoDB (2025). "MongoDB Acquires Voyage AI." https://www.mongodb.com/company/newsroom/press-releases/mongodb-acquires-voyage-ai-to-power-more-accurate-and-trustworthy-ai-applications
- Vectara (2025). "Hallucination Leaderboard." https://github.com/vectara/hallucination-leaderboard
- Dahl, M., Magesh, V., Suzgun, M., and Ho, D. E. (2024). "Large Legal Fictions." https://hai.stanford.edu/news/hallucinating-law-legal-mistakes-large-language-models-are-pervasive
- OpenAI (2023). "New models and developer products announced at DevDay." https://openai.com/index/new-models-and-developer-products-announced-at-devday/
- Regulation (EU) 2024/1689. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- NIST (2023). "Artificial Intelligence Risk Management Framework." https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
- LangChain (2024). "State of AI Agents." https://blog.langchain.dev/state-of-ai-agents/