Sai Tarun
Case study

Clinical RAG platform

Grounded answers over 50,000+ clinical and policy documents for healthcare authorization teams, with every release gated by evaluations.

Role
AI Software Engineer, contract
Company
Molina Healthcare
Timeline
Jun 2025 to Mar 2026
Stack
FastAPI, LangGraph, Pinecone, pgvector

The problem

Authorization reviewers had to dig through thousands of clinical and policy documents to find the criteria behind each request. It was slow, and in healthcare a confident wrong answer is worse than no answer at all.

The idea

Treat retrieval, grounding and evaluation as the product, not the LLM call.

  1. Documents
  2. Chunk and tag
  3. Embed
  4. Pinecone, pgvector
  1. Retrieve
  2. OpenAI, Claude
  3. Grounding check
  4. Answer with sources
  • Chunks follow document structure, not fixed lengths.
  • Metadata like document type and effective date narrows every search.
  • Answers must be supported by the retrieved evidence.

What I built

Ingestion
Parsing, semantic chunking and metadata enrichment for 50K+ clinical and policy documents, then embeddings into Pinecone and pgvector.
Services
FastAPI microservices for document processing, retrieval and LLM inference, reused across 4 authorization and policy workflows by 500+ people.
Agents
LangGraph workflows that plan, call retrieval tools, analyze evidence and summarize authorization packets, built with clinical SMEs and product managers.
Generation
OpenAI and Claude behind one interface, with function calling and structured outputs instead of free-form text.
Grounding
Validation logic that checks generated answers against the retrieved evidence.

Safety for HIPAA workflows

  • PHI redactionProtected health information is redacted before it reaches downstream processing and logs.
  • Token-based authEvery service call is authenticated.
  • Audit loggingWho did what, when, in which workflow.
  • Release gatesRAGAS and LangSmith evaluations over 200+ curated cases had to pass before every release.

Impact

Each number compares the new system with the previous manual workflow on the same set of tasks.

  • 40%faster document lookupMedian time to find the evidence needed for an authorization question.
  • 35%less manual reviewReviewer time per authorization packet.
  • 30%lower hallucination rateRelative drop in answer claims not supported by retrieved evidence, across 200+ curated cases.
  • WeeklyreleasesDown from every two weeks, with Docker and GitHub Actions CI/CD.

The hardest part

Retrieval quality, not the modelIf the wrong policy section came back, even a strong model produced a bad answer. Most of the work went into chunking, metadata, retrieval filters and the evaluation set, so every change could be measured instead of guessed.