AI web grounding: how to ground your LLMs and agents in live web data
Large language models are frozen at their training cutoff, so they hallucinate on anything recent or specific. Web grounding fixes that by feeding a model live, cited web content at query time.
This guide covers what web grounding is, how it differs from RAG, and how to add real-time web grounding to your LLMs and AI agents with a single API call to the Brave Search API. You don’t need a scraping, cleaning, or chunking pipeline.
What is AI web grounding? AI web grounding is the practice of supplying a language model with live, retrieved web content at query time, so its answers reflect current, verifiable sources instead of only its training data. It’s how a model moves from guessing to citing.
What’s the difference between web grounding and RAG?
Retrieval-augmented generation (RAG) is an architecture: retrieve relevant documents, then let the model generate an answer from them. Grounding is the goal: keeping a model’s output anchored to real, verifiable facts. RAG is one tactic for getting there, and it’s a good one for static, internal knowledge like your own docs and support tickets.
The problem is that RAG for internal documents can’t answer questions about the live world, such as today’s release notes, this morning’s incident, or last quarter’s numbers. For that, grounding has to reach the open web in real time. So the useful way to frame it is:
- RAG is the tactic; real-time web grounding is the goal.
A static corpus of data can handle what your company already knows. A web grounding step handles everything that changed after your model (and your vector store) was last updated.
Why real-time web grounding matters
Two failure modes push teams toward web grounding:
- Data staleness: a model’s weights are a snapshot, so it’s confidently wrong about anything after its cutoff.
- Hallucination: with no source in front of it, a model fills gaps with plausible fiction.
Grounding a model in live web results attacks both of these failures at once. It injects real-time context and gives the model something true to stand on, which helps reduce hallucinations.
For technical decision-makers, there’s a third driver: citations and auditability. Grounded answers can carry source URLs and titles, so a claim can be traced back and verified. In regulated or high-stakes settings, that audit trail is often the difference between a proof of concept and a shippable product.
The hidden cost of a do-it-yourself grounding pipeline
Most teams build web grounding the hard way, with a classic pipeline that glues together four steps:
- Call a search API to get links.
- Launch a headless browser (such as Playwright or Puppeteer) to fetch each page.
- Run a custom HTML parser and regex to clean the markup.
- Chunk and embed what’s left.
Each of these steps adds latency, cost, and another piece of code to maintain.
It’s also wasteful for an LLM. Raw pages are full of navigation bars, ads, and boilerplate that bloat your token count and your inference bill without adding signal. What a model actually needs is the substantive text, tables, and code from a page, already extracted. The Brave Search API collapses that step into a single call.
Web grounding in one API call: the LLM Context endpoint
Brave’s LLM Context endpoint is built for grounding. Instead of returning ten blue links for a human, it searches Brave’s index, extracts the relevant content, and returns pre-chunked, relevance-ranked text, tables, and code that are ready to drop into a prompt. You don’t need scraping, HTML cleaning, or a separate extraction service.
You also get inline controls that keep a live web RAG pipeline fast and cheap:
maximum_number_of_tokens(range 1024–32768) caps how much context comes back, so your prompt stays within budget.context_threshold_mode(strict,balanced,lenient) filters out low-relevance noise.- Goggles let you re-rank the web inline, boosting authoritative domains or discarding spam, with no post-processing.
python
import requests
# Brave's LLM Context endpoint returns ready-to-use, pre-extracted web content
url = "https://api.search.brave.com/res/v1/llm/context"
headers = {
"Accept": "application/json",
"X-Subscription-Token": "YOUR_BRAVE_API_KEY",
}
params = {
"q": "How to implement asyncio in Python 3.12",
"maximum_number_of_tokens": 4096, # keep context (and inference cost) bounded
"context_threshold_mode": "strict", # drop low-relevance content automatically
# Inline Goggle: drop noisy sources so only substantive pages ground the model
"goggles": "$discard,site=pinterest.com\n$discard,site=quora.com",
}
resp = requests.get(url, headers=headers, params=params)
data = resp.json()
# `grounding.generic` holds the extracted snippets per URL — join them into your prompt.
context = "\n\n".join(
f"{item['title']} ({item['url']}):\n" + "\n".join(item["snippets"])
for item in data["grounding"]["generic"]
)
# `sources` is keyed by URL — use it to attach citations to your answer.
for source_url, meta in data["sources"].items():
print(source_url, "—", meta.get("title"))The response already separates the grounding content from a sources map keyed by URL, so wiring up citations is trivial. You cite straight from sources with no extra work.
Drop-in grounded answers with the OpenAI SDK
If you’d rather have Brave write the grounded answer for you, citations included, the Answers endpoint is OpenAI-compatible. You don’t need to learn a new framework. You point the standard OpenAI client at Brave’s base URL and change the model name, and your existing chatbot has real-time web access.
python
from openai import OpenAI
# Point the standard OpenAI client at Brave's OpenAI-compatible endpoint
client = OpenAI(
api_key="YOUR_BRAVE_API_KEY",
base_url="https://api.search.brave.com/res/v1",
)
stream = client.chat.completions.create(
model="brave", # Brave's grounded-answers model
messages=[
{"role": "user", "content": "What were the major updates in the latest PyTorch release?"}
],
stream=True, # streaming is required for citations
extra_body={"enable_citations": True}, # return verifiable inline citations
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")There are two additional details worth knowing:
- Citations (and research mode) require
stream=True. - Brave returns richer data than a generalized completion, so citations arrive in the streamed payload alongside the answer text. They stream as tagged JSON inside the content deltas, and the sample above prints them raw. See the Answers docs for a complete example that parses them into numbered links.
Traditional RAG stack vs. Brave grounding
| Task | Traditional grounding stack | Brave Search API |
|---|---|---|
| Search | Third-party engine API or scraper | Independent first-party index |
| Extraction | Headless browser (Puppeteer / Playwright) | Built-in smart chunking (/llm/context) |
| Cleaning | Custom HTML parser and regex | Pre-formatted text, tables, and code |
| Re-ranking | Vector database + embedding model | Inline Goggles + threshold modes |
| Latency | Higher (several services in series) | Low (single API call) |
Grounding for AI agents, MCP, and your framework
Web grounding is really a tool call: an agent decides it needs fresh facts, calls a search endpoint, and reasons over what comes back. Brave fits that shape directly. The official Brave Search MCP server lets any Model Context Protocol client give its agent web grounding as a native tool. For agentic frameworks, there are ready-made integrations for LangChain and LlamaIndex.
Grounding shouldn’t require adopting a whole platform. The Brave Search API is a single-purpose primitive (a building block) that plugs into the model and framework you already use, whether that’s MCP, the OpenAI SDK, LangChain, LlamaIndex, or something else. That’s simpler than tying your grounding layer to a single model vendor’s built-in search tool or cloud platform. With Brave, you keep your model, your framework, and your stack.
Why the index matters: independence, privacy, and citations
Not every “grounding” source is equal, because most don’t run their own search. Many providers resell or scrape another engine’s results. That caps quality and ties your product to a competitor’s rate limits and terms. Brave grounds on its own full-scale, independent web index of 40+ billion pages, one of the few outside Big Tech at this scale. You get results built for machines, not a repackaged feed of someone else’s ten blue links.
Independence also unlocks two things enterprises care about:
- Privacy: Brave owns its entire search stack, from crawler to API endpoint, and doesn’t scrape other engines. That lets it offer true, architectural Zero Data Retention. Enterprise customers can enable ZDR so no queries are retained for any length of time. On standard plans, query records are kept for a maximum of 90 days and used for limited purposes such as billing, troubleshooting, and abuse prevention.
- Authenticity: some engines push publishers toward “Generative Engine Optimization,” reformatting the web to feed a specific AI. Brave indexes the web as it actually is. Structured source URLs and titles come out of the box, so citation enforcement and audit trails become the default, not a project.
Start now
Grounding your AI in live web data takes one API call and a few minutes. Get a key and try the LLM Context and Answers endpoints on your own queries. The Search and Answers plans include $5 in free monthly credits to start.
→ Get your Brave Search API key
Frequently asked questions
What is web grounding for AI? Web grounding is supplying an AI model with live web content at query time so its answers reflect current, verifiable sources. It reduces hallucinations and fixes the staleness that comes from a fixed training cutoff.
What’s the difference between grounding and RAG? RAG is an architectural pattern (retrieve, then generate). Grounding is the goal of keeping output anchored to real facts. Document RAG grounds a model in your static internal files; real-time web grounding grounds it in the live open web.
Does web grounding reduce AI hallucinations? Yes. Giving a model retrieved, cited sources to reason over reduces hallucinations, because it no longer has to fill gaps from memory. It doesn’t eliminate them entirely, so citations remain important for verification.
Can I add web grounding without changing my stack? Yes. Brave’s Answers endpoint is OpenAI-SDK-compatible: swap the base URL and use model="brave". There are also MCP, LangChain, and LlamaIndex integrations, so you can add grounding to your existing framework in minutes.
Is the Brave Search API private? Brave serves results from its own independent index rather than reselling another engine’s results. On standard plans, query records are kept for a maximum of 90 days for purposes like billing, troubleshooting, and abuse prevention. For compliance-sensitive RAG applications, Enterprise customers can enable true Zero Data Retention, so queries aren’t retained at all.