Saltar al contenido principal

Una publicación etiquetados con "embeddings"

Ver Todas las Etiquetas
advancedPart 5

The capstone: build, evaluate, and observe a RAG docs assistant on the WEC API

· 27 min de lectura
Rafael Fernandes
Ingeniero de PLN y Redactor Técnico en WiLine
Share:
+Promptfoo+Langfuse
0/6
🎯 Skill path0/6 earned
AI evals & observability

Everything in this series was building to this. You can prove a model works (part 1), make its output machine-reliable (part 2), generate real test data (part 3), and observe production (part 4). Now we spend all four skills at once on the pattern behind almost every serious LLM product: RAG — retrieval-augmented generation.

We'll build a docs assistant: a containerized HTTP service that answers customer questions from WEC's own documentation. Not a notebook — a service, Docker-first, the shape you'd actually deploy. And because this series doesn't do happy-path demos: along the way our RAG hallucinates a GPU price, we root-cause it to our own scraper, fix it, and pin the fix with a regression test. Every command, number, and error below is from a real run.