Curated example · Editable connections
RAG architecture
you can explore.
Separate document ingestion from retrieval and answer generation.
Curated example
Follow the connections.
Source material with ownership, freshness, and access rules.
Make it your own
Documents -> Chunking -> Document Embeddings -> Vector Store
Question -> Query Embedding -> Retrieval
Vector Store -> Retrieval -> Context Assembly -> Language Model -> Answer
Question -> Context AssemblyChecking workspace availability. You can explore and copy this example now.
Two paths, one shared index
The ingestion path prepares documents for retrieval. The serving path embeds a question, retrieves relevant passages, and passes context to a model. Showing the paths separately makes it easier to discuss updates, latency, and where data is allowed to move.
Extend the RAG pipeline diagram
Add a reranker between retrieval and context assembly if your system uses one. Add an evaluation step outside the main serving path to inspect retrieval quality and answer grounding. Model names and vector databases are implementation choices; the responsibilities remain useful to show.
Make boundaries explicit
Identify which documents a user may access before assembling context. Mark stale or deleted sources and decide how the index is refreshed. This is a curated design example, not a deployed RAG application or evidence that a model will answer correctly.
Further reading: Google Cloud: RAG reference architecture ↗. The example above is independently authored and simplified.
Another system to explore
SaaS architecture
Follow a web request through an API, a cache, and a database.
Explore the diagram02AI agent architecture
See how an orchestrator, model, tools, and memory work together.
Explore the diagram03RAG architecture
Separate document ingestion from retrieval and answer generation.
Explore the diagram