Clinical RAG platform
Grounded answers over 50,000+ clinical and policy documents for healthcare authorization teams, with every release gated by evaluations.
- Role
- AI Software Engineer, contract
- Company
- Molina Healthcare
- Timeline
- Jun 2025 to Mar 2026
- Stack
- FastAPI, LangGraph, Pinecone, pgvector
The problem
Authorization reviewers had to dig through thousands of clinical and policy documents to find the criteria behind each request. It was slow, and in healthcare a confident wrong answer is worse than no answer at all.
The idea
Treat retrieval, grounding and evaluation as the product, not the LLM call.
- Documents
- Chunk and tag
- Embed
- Pinecone, pgvector
- Retrieve
- OpenAI, Claude
- Grounding check
- Answer with sources
- Chunks follow document structure, not fixed lengths.
- Metadata like document type and effective date narrows every search.
- Answers must be supported by the retrieved evidence.
What I built
- Ingestion
- Parsing, semantic chunking and metadata enrichment for 50K+ clinical and policy documents, then embeddings into Pinecone and pgvector.
- Services
- FastAPI microservices for document processing, retrieval and LLM inference, reused across 4 authorization and policy workflows by 500+ people.
- Agents
- LangGraph workflows that plan, call retrieval tools, analyze evidence and summarize authorization packets, built with clinical SMEs and product managers.
- Generation
- OpenAI and Claude behind one interface, with function calling and structured outputs instead of free-form text.
- Grounding
- Validation logic that checks generated answers against the retrieved evidence.
Safety for HIPAA workflows
- PHI redactionProtected health information is redacted before it reaches downstream processing and logs.
- Token-based authEvery service call is authenticated.
- Audit loggingWho did what, when, in which workflow.
- Release gatesRAGAS and LangSmith evaluations over 200+ curated cases had to pass before every release.
Impact
Each number compares the new system with the previous manual workflow on the same set of tasks.
- 40%faster document lookupMedian time to find the evidence needed for an authorization question.
- 35%less manual reviewReviewer time per authorization packet.
- 30%lower hallucination rateRelative drop in answer claims not supported by retrieved evidence, across 200+ curated cases.
- WeeklyreleasesDown from every two weeks, with Docker and GitHub Actions CI/CD.
The hardest part
Retrieval quality, not the modelIf the wrong policy section came back, even a strong model produced a bad answer. Most of the work went into chunking, metadata, retrieval filters and the evaluation set, so every change could be measured instead of guessed.