Teaching an AI to build software: the case for platform-generation engines

Why the next leap in platform engineering is giving AI the keys to the entire SDLC

by
Teaching an AI to build software: the case for platform-generation engines

Teaching an AI to build software: the case for platform-generation engines

We’ve all seen the demos: an LLM spits out a React component or a Python script from a prompt. Impressive, but useless for production. The real prize isn’t generating snippets—it’s teaching an AI to own the entire software development lifecycle, from a vague requirement to a deployed, tested, and monitored system.

This is what I call a platform-generation engine (PGE). It’s an autonomous agent—or a swarm of agents—that takes a high-level goal and produces a working software platform. No human writes a single line of code. The AI builds the CI/CD pipeline, the database schema, the API contracts, the tests, and the monitoring dashboards. It self-corrects when something breaks. It refactors when requirements change.

This sounds like science fiction, but the pieces exist today. The missing link is architecture: how to wire together LLMs, sandboxed execution, vector databases, and feedback loops into a system that can be trusted to build production software.

Why platform engineering is the perfect use case

Platform engineering is already about abstraction—building internal developer platforms that hide complexity. A PGE takes that abstraction to the next level: the platform itself becomes a product of AI generation.

Consider the typical platform engineering workflow:

  1. Gather requirements from stakeholders.
  2. Design architecture.
  3. Scaffold repositories, CI/CD, infrastructure-as-code.
  4. Implement services, databases, APIs.
  5. Write tests, documentation, runbooks.
  6. Deploy, monitor, iterate.

A PGE automates steps 2 through 6. The human remains in the loop for step 1 (and for approving critical changes), but the heavy lifting is done by agents.

This is not the same as low-code or no-code platforms. Those are rigid—you can only build what the platform allows. A PGE uses LLMs to generate arbitrary code, configuration, and infrastructure. It’s constrained only by the capabilities of the underlying models and the sandbox environment.

The architecture of a platform-generation engine

I’ve been building a PGE called GenForge (working name, open-source, not yet released). The architecture is straightforward:

1. Orchestrator agent

A top-level agent that receives the goal (e.g., “Build a multi-tenant SaaS app for project management with PostgreSQL, Redis, and a React frontend”). It decomposes this into sub-tasks: schema design, API design, frontend scaffolding, CI/CD setup, etc. Each sub-task is assigned to a specialized agent.

We use LangGraph for orchestration with a GPT-4o or Claude 3.5 Sonnet model. The orchestrator maintains a shared state in PostgreSQL with pgvector for storing intermediate artifacts and their embeddings for retrieval.

2. Specialized agents

Each agent is a loop:

  • Code generator: Takes a spec, generates code files. We use vLLM serving a fine-tuned CodeLlama-34B (LoRA adapters for specific frameworks). The agent outputs a directory structure.
  • Tester: Runs the generated code in a sandbox (Docker container with network restrictions). Executes unit tests, integration tests, linting. If tests fail, it sends the error back to the code generator with a request to fix.
  • Infrastructure agent: Generates Terraform or Pulumi code for cloud resources. It has access to a Qdrant vector store of typical infrastructure patterns (VPC, RDS, EKS, etc.).
  • Documentation agent: Generates README, API docs (OpenAPI), and runbooks from the code and configuration.

All agents log their actions to a central PostgreSQL database (with pgvector for similarity search on past failures/solutions).

3. Sandboxed execution environment

This is critical. The PGE must be able to execute generated code without blowing up production. We use Firecracker microVMs (via flyte or custom orchestration) for each generation run. Each VM has a base image with common runtimes (Python, Node, Go) and is wiped after the run. Network access is limited to package registries (PyPI, npm) and nothing else.

4. Feedback loop

After the initial generation, the PGE deploys to a staging environment (a separate Kubernetes namespace). It runs a battery of synthetic tests (using k6 for load, Selenium for UI). If the system fails any test, the orchestrator creates a new sub-task to fix the issue, and the cycle repeats.

This feedback loop is what makes the PGE self-correcting. Without it, you’re just generating code and hoping it works.

A concrete example: building a microservice

Let’s walk through a real (simplified) run of GenForge. The goal: “Build a user management microservice with CRUD endpoints, PostgreSQL persistence, and JWT authentication.”

  1. Orchestrator parses the goal and creates sub-tasks:

    • Design database schema (users table, roles, etc.)
    • Generate FastAPI app with routes
    • Generate migration scripts (Alembic)
    • Generate Dockerfile and docker-compose
    • Generate tests (pytest)
    • Generate CI/CD (GitHub Actions)
  2. Schema agent outputs:

    CREATE TABLE users (
        id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
        email VARCHAR(255) UNIQUE NOT NULL,
        password_hash VARCHAR(255) NOT NULL,
        role VARCHAR(50) DEFAULT 'user',
        created_at TIMESTAMPTZ DEFAULT NOW()
    );

    It also generates the Alembic migration.

  3. Code agent generates:

    # app/main.py
    from fastapi import FastAPI, Depends, HTTPException
    from sqlalchemy.orm import Session
    from . import models, schemas, auth
    from .database import SessionLocal, engine
    
    models.Base.metadata.create_all(bind=engine)
    
    app = FastAPI()
    
    @app.post("/users/", response_model=schemas.User)
    def create_user(user: schemas.UserCreate, db: Session = Depends(get_db)):
        # ...
  4. Tester agent runs pytest inside the sandbox. It finds a test failure because the password hashing isn’t implemented. It sends the error back to the code agent.

  5. Code agent fixes the issue by adding auth.py with passlib hashing.

  6. Infrastructure agent generates a docker-compose.yml with the app and PostgreSQL.

  7. CI/CD agent generates a GitHub Actions workflow that builds, tests, and deploys to a staging server.

  8. Orchestrator deploys to staging, runs k6 load tests (100 concurrent users). If the service handles it, it marks the task as done.

All of this happens without human intervention. The entire process took about 12 minutes on my test setup (using GPT-4o for the orchestrator and CodeLlama for code generation).

Data sovereignty and self-hosting

For a PGE to be viable in regulated environments, it must be self-hosted. You cannot send your company’s source code or architecture to OpenAI’s servers. That’s why we built GenForge to run entirely on-prem.

  • LLM inference: We use vLLM with Llama 3.1 70B (quantized to 4-bit) for the orchestrator, and CodeLlama-34B with LoRA adapters for code generation. Both run on a single node with 4x A100 80GB GPUs.
  • Vector store: Qdrant (or pgvector for simplicity) for storing past solutions and patterns.
  • Execution sandbox: Firecracker microVMs managed by a custom Go service.
  • Orchestration: LangGraph running in a Python process behind systemd.

Everything is containerized with Docker and orchestrated with docker-compose for simplicity (or Kubernetes for scale).

The challenges

This is not a solved problem. Here are the biggest hurdles:

1. Hallucination in infrastructure code

LLMs are great at generating Terraform, but they often hallucinate resource names, regions, or IAM policies. The only mitigation is to run terraform plan in the sandbox and parse the output. If the plan has errors, the agent retries with the error message.

We’ve built a small Terraform validator agent that checks the generated HCL against a set of company policies (e.g., “no public S3 buckets”). This is a RAG pipeline using BGE-M3 embeddings and Qdrant.

2. Long context windows

A typical PGE run generates thousands of lines of code across dozens of files. The orchestrator’s context window fills up quickly. We use sliding window summarization: after each sub-task completes, the orchestrator summarizes the output and stores the full artifact in PostgreSQL. The summary is kept in context; the full artifact is retrieved only when needed via vector search.

3. Debugging agent loops

When an agent gets stuck in a loop (e.g., code gen -> test fail -> code gen -> test fail ad infinitum), you need a circuit breaker. We use a retry limit (3 per sub-task) and a human-in-the-loop hook: if the limit is exceeded, the orchestrator pauses and asks a human for guidance. The human’s input is stored in the vector store for future reference.

Is this ready for production?

Not yet. But it’s close. I’ve used GenForge to generate internal tools (a simple inventory management app, a Slack bot for incident response) that are running in production today. The key is to start small: let the PGE generate only the scaffolding and boilerplate, and have human developers fill in the business logic. Over time, as the feedback loop improves, you can increase the scope.

For platform engineering teams, a PGE is not a replacement—it’s a force multiplier. It handles the drudgery of setting up repos, CI/CD, and basic services, freeing humans to focus on architecture, security, and complex business rules.

The road ahead

We’re releasing GenForge as open-source next month (watch the repo at github.com/dradulic/genforge). The initial release will support Python/Node services, PostgreSQL, Docker, and GitHub Actions. Future versions will add support for event-driven architectures, message queues, and more complex deployment targets (Kubernetes, serverless).

If you’re building a platform team, I’d encourage you to experiment with PGEs. Start by giving an LLM agent a simple spec for a microservice and see how far it gets. You’ll be surprised—and you’ll see exactly where the gaps are. That’s the first step toward closing them.

#ai-agents#autonomy#codegen#platform-engineering#self-hosted
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.