Why a Swarm of Small Agents Beats a Giant Model for Civic Intelligence

Concrete benchmarks from rinet.one

by
Why a Swarm of Small Agents Beats a Giant Model for Civic Intelligence

Why a Swarm of Small Agents Beats a Giant Model for Civic Intelligence

Conventional wisdom says bigger models are better. For civic intelligence tasks — analyzing legislation, tracking public sentiment, monitoring policy changes — we found the opposite. A swarm of small, specialized agents consistently outperformed a single monolithic LLM in accuracy, cost, and latency.

This isn't a thought experiment. We built it. We benchmarked it. Here are the numbers.

The Problem: Civic Intelligence at Scale

Civic intelligence means extracting actionable signals from messy public data: city council minutes, regulatory filings, news articles, social media. The data is heterogeneous, noisy, and time-sensitive. A single model must be good at everything: summarization, entity extraction, sentiment analysis, cross-referencing, and more.

A monolithic model like GPT-4 or Llama-3-70B can do all of these, but it's expensive and slow. More importantly, it's brittle: one prompt change can break performance on a subtask. And if you need to update one capability, you retrain the whole model.

The Swarm Architecture

We built a multi-agent system called rinet.one that decomposes civic intelligence into specialized subtasks. Each subtask is handled by a small, fine-tuned model (typically Llama-3-8B or Mistral-7B) or a lightweight RAG pipeline. The agents communicate via a shared Postgres-backed message bus.

Agent Roles

  • Scraper Agent: Fetches documents from specified sources (e.g., gov websites, RSS feeds). Uses a lightweight headless browser (Playwright) and stores raw HTML in Postgres.
  • Parser Agent: Converts HTML to structured text using a fine-tuned BERT-like model for layout detection. Outputs markdown with metadata.
  • Extractor Agent: Extracts entities (people, organizations, dates, locations) using a small NER model (e.g., GLiNER with BGE-M3 embeddings).
  • Summarizer Agent: Generates concise summaries using a 7B parameter model quantized to 4-bit (llama.cpp with Q4_K_M).
  • Sentiment Agent: Classifies sentiment on a 5-point scale using a distilled RoBERTa model.
  • Cross-Ref Agent: Links entities across documents using vector similarity in pgvector (embedding dim 1024 from BGE-M3).
  • Alert Agent: Checks for predefined patterns (e.g., budget cuts, policy changes) using a rule-based system with regex and small classifiers.

Each agent runs as a systemd service, communicating via JSON messages over a Redis queue. The supervisor agent orchestrates the pipeline, handling retries and dead-letter queues.

Benchmark Setup

We compared the swarm against two monolithic baselines:

  • GPT-4o: The current state-of-the-art from OpenAI.
  • Llama-3-70B: Self-hosted on 2x A100 GPUs using vLLM.

Task Suite

We created a benchmark of 500 civic intelligence queries covering three categories:

  1. Entity Extraction: Extract all organizations, people, and dates from a city council meeting transcript.
  2. Cross-Document Linking: Given two documents about the same policy, identify the common entities and summarize the differences.
  3. Sentiment & Alert: Classify the sentiment of a public comment and trigger an alert if it mentions a specific budget item.

Each query had a ground-truth answer verified by human annotators.

Metrics

  • Accuracy: Exact match for entities, F1 for linking, 5-point scale agreement for sentiment.
  • Cost: API costs for GPT-4o, estimated GPU cost for Llama-3-70B (at $1.50/hr per A100), and total compute cost for the swarm (CPU + GPU).
  • Latency: End-to-end time from query submission to final output.

Results

Accuracy

Task GPT-4o Llama-3-70B Swarm
Entity Extraction 0.89 0.87 0.94
Cross-Document Linking 0.82 0.79 0.91
Sentiment & Alert 0.91 0.88 0.95

Swarm outperformed on all tasks. The biggest gap was in cross-document linking, where the dedicated Cross-Ref Agent with pgvector similarity search beat the monolithic models' context-limited reasoning.

Cost per Query

Model Cost per Query
GPT-4o $0.15
Llama-3-70B $0.08
Swarm $0.02

Swarm costs 75% less than Llama-3-70B and 87% less than GPT-4o. The savings come from using smaller models and CPU-only agents for most tasks. Only the Summarizer Agent needs GPU, and it's a 7B quantized model.

Latency (P50, seconds)

Model Latency
GPT-4o 12.3
Llama-3-70B 8.7
Swarm 3.1

Swarm is 2.8x faster than Llama-3-70B and 4x faster than GPT-4o. The parallel execution of agents (Scraper, Parser, Extractor run concurrently) reduces wall-clock time.

Why the Swarm Wins

Specialization beats generalization

Each small agent is fine-tuned for exactly one task. The Extractor Agent uses a model trained only on legal and governmental text. The Sentiment Agent sees only public comments. This focus yields higher accuracy than a generalist model trying to do everything.

Modular updates

When a new document format appears (e.g., a city switches from PDF to HTML), we update only the Parser Agent. No retraining of the whole system. In production, we've updated agents independently without downtime.

Cost-efficient scaling

Adding a new civic source (e.g., a county board) doesn't require more GPU power. We spin up a new Scraper Agent and route its output through existing agents. The bottleneck is database throughput, not model inference.

Robustness through redundancy

If one agent fails (e.g., the GPU node for Summarizer goes down), the swarm degrades gracefully. The Alert Agent can still run on CPU-only rules. Monolithic models fail entirely if the GPU is unavailable.

Implementation Details

Postgres as Control Plane

All agent state, intermediate results, and final outputs are stored in Postgres 16 with pgvector 0.7.0. We use Postgres as the single source of truth: agents read tasks from a jobs table, write results to outputs, and report status to agent_heartbeat. This gives us full auditability and easy debugging.

Embeddings with BGE-M3

For cross-document linking, we use BGE-M3 (BAAI/bge-m3) with 1024-dimensional embeddings. Documents are chunked into 512-token windows with 128-token overlap. Embeddings are indexed with IVFFlat (lists=100) in pgvector for fast similarity search.

Model Serving

  • Summarizer: llama.cpp server with Llama-3-8B-Instruct quantized to Q4_K_M. Runs on a single RTX 4090. Context window of 8192 tokens.
  • Extractor: GLiNER model (urchade/gliner_multi-v2.1) with BGE-M3 embeddings for entity recognition. CPU-only, runs on the same machine as Postgres.
  • Sentiment: DistilRoBERTa-base fine-tuned on civic sentiment data. CPU-only, 100ms per inference.

Orchestration

Agents communicate via Redis Streams. Each agent subscribes to an input channel and publishes to an output channel. The Supervisor Agent (a Python script using asyncio) monitors job progress and handles retries with exponential backoff (max 3 retries). Dead jobs go to a failed_jobs table for manual inspection.

Case Study: Tracking a School Board Budget

We deployed the swarm to track a large school district's budget negotiations. The district publishes weekly meeting minutes, budget spreadsheets, and public comments. The swarm processed 2,000 documents over 3 months.

Results:

  • Extracted 1,500+ entities (board members, schools, budget line items).
  • Linked 80% of public comments to specific budget items.
  • Triggered 12 alerts when the board discussed teacher salaries.
  • Total cost: $120 in compute (vs. estimated $900 for GPT-4o).

When to Use a Monolithic Model

Swarm architecture isn't always better. If your task is a single, well-defined generation (e.g., write a poem), a monolithic model is simpler. But for complex, multi-step reasoning over diverse data sources, the swarm wins on every axis.

The Future

We're working on dynamic agent creation: when the system encounters a novel task type, it spawns a new agent fine-tuned on the fly using LoRA adapters. The base model remains frozen; only the adapter weights are updated. This brings the cost of adding a new capability to under $10 in compute.

Build Your Own

You don't need a massive cluster. A single machine with a consumer GPU (RTX 4090, 24GB VRAM) and 64GB RAM can run the entire swarm. Start with the reference implementation at github.com/rinetone/swarm.

Monolithic models are yesterday's solution. For civic intelligence, the future is a swarm.

#agentic#benchmarks#multi-agent#rag
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related