Research agent with its own benchmark
A RAG service, a four-agent research loop and a benchmark that tests whether the loop beats a single LLM call.
Personal projectAuthor2026
Problem
Multi-agent systems are easy to demo. I wanted to know whether a multi-agent loop actually beats a single LLM call on the same questions.
Solution
Three parts that build on each other: ai-document-analyst, a RAG service with hybrid retrieval; research-agent, a four-role agent that uses it as a tool; and agent-benchmark, a harness that runs both approaches on the same questions.
Architecture
The agent loop
- Question
- PlannerPlans the research
- ResearcherCalls the RAG service
- CriticReviews the draft
- SynthesizerWrites the final answer
Inside the RAG service
- BM25Keyword match
- Dense vectorsSemantic match
- Cross-encoderReranks the merged results
- The benchmark harness runs the agent loop and a single-shot baseline on the same questions.
- CI and tests run without API keys, so anyone can reproduce the setup.
Stack
PythonFastAPIReactDockerGitHub Actionspytest
Result
A reproducible way to compare the agent loop with a single LLM call, instead of judging it by a demo.