Juan Pablo García
All projects

Research agent with its own benchmark

A RAG service, a four-agent research loop and a benchmark that tests whether the loop beats a single LLM call.

Personal projectAuthor2026

Problem

Multi-agent systems are easy to demo. I wanted to know whether a multi-agent loop actually beats a single LLM call on the same questions.

Solution

Three parts that build on each other: ai-document-analyst, a RAG service with hybrid retrieval; research-agent, a four-role agent that uses it as a tool; and agent-benchmark, a harness that runs both approaches on the same questions.

Architecture

The agent loop

  1. Question
  2. PlannerPlans the research
  3. ResearcherCalls the RAG service
  4. CriticReviews the draft
  5. SynthesizerWrites the final answer

Inside the RAG service

  1. BM25Keyword match
  2. Dense vectorsSemantic match
  3. Cross-encoderReranks the merged results
  • The benchmark harness runs the agent loop and a single-shot baseline on the same questions.
  • CI and tests run without API keys, so anyone can reproduce the setup.

Stack

PythonFastAPIReactDockerGitHub Actionspytest

Result

A reproducible way to compare the agent loop with a single LLM call, instead of judging it by a demo.

Have an AI feature to build? Let's talk for 15 minutes.