From Chatbot to Co-Founder: Running an AI That Builds AI
A technical case study in autonomous agent swarms that fine-tune their own models and evolve their own stack.
Most people treat LLMs as chatbots. You type, it answers. You get the same model tomorrow as you did today. That's fine for customer support. It's useless for building a company.
We wanted something different: an AI that doesn't just talk but acts. An AI that improves itself, writes its own code, fine-tunes its own models, and manages its own infrastructure. An AI that acts less like a tool and more like a co-founder.
This is the story of how we built that system. No hype. Just the architecture, the tools, the failures, and the lessons learned.
The Vision: An AI That Grows
The core idea is simple: an agent should not be static. It should learn from every interaction, every mistake, every success. It should fine-tune its own base models, evolve its prompts, and even rewrite its own orchestration code.
We call this a "self-evolving agent swarm." It's a multi-agent system where one agent is dedicated to improving the others. It monitors performance, collects training data, launches fine-tuning jobs, and deploys updated models.
Architecture Overview
Our stack runs entirely on-prem. No cloud APIs for the core inference. Data sovereignty was a hard requirement — we wanted full control over our models and data.
Core components:
- vLLM for high-throughput LLM inference. We run multiple instances of open-source 7B and 8B parameter models.
- llama.cpp for edge cases where latency matters more than throughput (single-user interactive debugging).
- PostgreSQL as the control plane. All agent state, task queues, and model metadata live in Postgres. pgvector for similarity search over past interactions.
- Qdrant for high-performance vector search on embeddings — used for retrieving few-shot examples during fine-tuning.
- BGE-M3 embedding model for converting agent logs into vector representations.
- Pgbouncer for connection pooling between agents and Postgres.
- Systemd for process supervision — each agent runs as a systemd service with automatic restart.
The Agent Swarm
We run three types of agents:
- Worker agents — handle user requests, execute tasks, generate code, answer questions.
- Supervisor agent — monitors worker performance, detects patterns of failure, and decides when to trigger a fine-tuning cycle.
- Builder agent — the most interesting one. It manages the fine-tuning pipeline and model deployment.
The Builder agent is what makes the system self-evolving. Here's how it works.
Self-Evolution Loop
Step 1: Collect Training Data
Every interaction between a worker agent and a user is logged into Postgres. The log includes:
- The raw prompt and response
- The user's explicit feedback (thumbs up/down)
- Implicit signals: did the user edit the output? Did they re-run the task? Did they ask a follow-up that suggests confusion?
A background process (a simple Go service) computes a quality score for each interaction. Interactions with high quality are marked as positive examples. Interactions with low quality are stored for negative sampling.
Step 2: Prepare Fine-Tuning Dataset
Once the Supervisor agent detects a pattern — say, the worker keeps failing on SQL generation tasks — it triggers the Builder agent.
The Builder agent queries Postgres for all relevant interactions (using pgvector to find semantically similar failures). It then formats them into a LoRA training dataset.
Step 3: Fine-Tune with LoRA
We use LoRA (Low-Rank Adaptation) because it's fast and cheap. We don't need to retrain the entire model.
The Builder agent spins up a fine-tuning job using a custom script that leverages Hugging Face's PEFT library. It runs on a dedicated GPU node.
Typical parameters:
- LoRA rank: 16
- Learning rate: 2e-4
- Batch size: 4
- Epochs: 3
- Target modules: q_proj, v_proj
The fine-tuning takes on the order of tens of minutes for a dataset of a thousand examples on a 7B model.
Step 4: Evaluate and Deploy
After fine-tuning, the Builder agent runs an evaluation harness. It tests the new model against a held-out set of tasks. If performance improves measurably (by task completion rate), the new LoRA adapter is deployed.
Deployment is simple: the Builder agent updates a symlink in the model directory and restarts the vLLM instance via systemd. The worker agents pick up the new model on the next request.
Step 5: Close the Loop
The Supervisor agent logs the outcome. If the fine-tuning didn't improve performance, it escalates to a human operator. So far, most fine-tuning cycles have been successful.
Code Evolution: The Agent That Rewrites Itself
Fine-tuning the LLM is one thing. But we also wanted the agent to improve its own code — the logic that governs its behavior.
We built a meta-agent that can read its own source code (Python, Go, and shell scripts), identify bottlenecks or bugs, and propose patches. The patches are tested in a sandboxed environment (Docker containers) before being merged.
Here's a simplified example of the meta-agent's workflow:
- It detects that the agent's prompt template is too verbose, causing high latency.
- It generates a shorter prompt template.
- It runs an A/B test: a fraction of traffic uses the new prompt, the rest uses the old.
- After a number of requests, it compares latency and success rate.
- If the new prompt is better, it updates the configuration file and restarts the agent.
All of this happens autonomously. No human in the loop.
Infrastructure as Code, Managed by AI
We didn't stop at code and models. The Builder agent also manages infrastructure. It monitors disk usage, GPU memory, and request queues. If it detects that the vLLM instance is overloaded, it can spin up a second instance on another GPU node.
It does this via Ansible playbooks that it generates on the fly. Yes, an AI is writing Ansible playbooks and applying them to production servers.
We have safety limits, of course. The agent cannot modify network rules or access secrets without human approval. But within those boundaries, it has full autonomy.
Data Sovereignty and the EU AI Act
Running this on-prem was not just a preference — it was a requirement. The EU AI Act classifies systems that use self-evolution as high-risk. We need to maintain audit logs for every model version and every training dataset.
Our Postgres database acts as the source of truth. Every fine-tuning run is logged with:
- The exact dataset used (referenced by a hash)
- The training parameters
- The evaluation results
- The deployment timestamp
This makes it easy to produce compliance reports. If a regulator asks "What did your model learn on March 15th?", we can answer precisely.
Lessons Learned
What Worked
- Postgres as the control plane was a game-changer. We tried Redis and etcd, but Postgres's relational model and pgvector support made it far easier to query historical data.
- LoRA fine-tuning is fast and effective for targeted improvements. We can fix specific failure modes without degrading general performance.
- Systemd for process management is boring but reliable. No Kubernetes overhead for a small cluster.
What Didn't Work
- Full fine-tuning was too slow and unstable. We tried fine-tuning the entire model once, and it took many hours on our hardware. The resulting model regressed on unrelated tasks. LoRA is the way to go.
- Letting the agent design its own fine-tuning architecture was a mistake. It once tried to use a rank that was too large, which consumed all VRAM and crashed the node. We now hardcode sensible ranges.
- Over-reliance on implicit feedback led to noisy training data. Users often don't edit output even when it's wrong. We now combine explicit and implicit signals with a confidence threshold.
Surprising Wins
- The Builder agent once discovered that a specific set of prompts caused the worker to hallucinate. It automatically generated a system prompt that reduced hallucinations substantially.
- The meta-agent refactored its own codebase to use async I/O, improving throughput significantly. We didn't tell it to do that.
The Bottom Line
We now run a system where the AI is genuinely self-improving. It fine-tunes its own models, rewrites its own code, and manages its own infrastructure. It's not perfect — we still have human oversight for critical decisions. But the trend is clear: the AI is becoming more autonomous over time.
This is not science fiction. It's a practical architecture built with open-source tools and commodity hardware. If you have a GPU, Postgres, and a willingness to experiment, you can build something similar.
The chatbot era is over. The co-founder era has begun.
References
- vLLM: https://github.com/vllm-project/vllm
- llama.cpp: https://github.com/ggerganov/llama.cpp
- PEFT/LoRA: https://github.com/huggingface/peft
- pgvector: https://github.com/pgvector/pgvector
- BGE-M3: https://huggingface.co/BAAI/bge-m3