AI Data Copilot
Ask a spreadsheet anything in plain English. The AI writes the analysis code, runs it safely, and repairs it when it breaks.
- Role
- Solo project
- Type
- Web app, live
- Model
- GPT-4.1 and 4.1-mini
- Stack
- Python, Streamlit, Pandas, Plotly
The problem
Most people with a spreadsheet have questions, not Python or SQL. Paste the data into a chatbot and you get an answer that sounds right, with no way to check where the number came from.
The idea
Let the model write code, not answers.
- Question
- Router
- GPT-4.1 writes pandas
- Sandbox
- Error? Retry, up to 3
- Chart and insights
- Every number comes from pandas running on the real data.
- The generated code and every attempt are shown to the user.
- Descriptive questions skip code entirely.
How it works
- Profiling
- The file is read, duplicate column names are cleaned up, and a profile of types, missing values and unique counts is built.
- Routing
- GPT-4.1-mini decides whether a question needs analysis or just a written answer, so simple questions stay fast and cheap.
- Code
- GPT-4.1 writes pandas code against the dataset profile, returning a result table and, when useful, a chart.
- Sandbox
- Code runs with a minimal set of Python builtins. Anything touching files, the network, the OS or imports is blocked before it runs.
- Self-repair
- If the code fails, the error goes back to the model as context for a fix, up to three times.
- Explaining
- GPT-4.1-mini turns the result into three insights, one recommendation and one caveat.
- Dashboards
- Schema detection finds date, revenue, profit, cost, discount, quantity and category columns, then builds a sales, classification or general dashboard.
- Reports
- Any analysis can be downloaded as a Markdown report, and every run is kept in a session history.
The hardest part
Letting an AI run code without letting it run anythingExecuting model-written code is the riskiest thing the app does. The answer was layers: screen every snippet before it runs, execute it with almost no builtins, and treat failures as information instead of crashes. A typo in a column name becomes a quiet retry with the error attached, not a stack trace in front of the user.
What I'd build next
- Isolation
- Move execution into a short-lived container per request. A blocklist is a good first layer, but not a true sandbox.
- Evaluation
- A set of questions with known answers, run on every change, so prompt edits cannot quietly break the numbers.
- Bigger data
- Multi-file and SQL sources, so it can work across related tables.