Building RiNET: A Sovereign Autonomous OS for Institutional Cognition

How we designed a decentralized, self-healing agent swarm system for enterprise AI autonomy

by
Building RiNET: A Sovereign Autonomous OS for Institutional Cognition

Institutional cognition is the collective intelligence of an organization — its knowledge, decision-making processes, and operational memory. Most organizations today rely on fragmented SaaS tools, brittle APIs, and centralized AI services that expose sensitive data to third parties. RiNET is our answer: a sovereign autonomous operating system that runs entirely on your infrastructure, orchestrates a swarm of fine-tuned agents, and self-heals without human intervention.

This article walks through the architecture, the engineering tradeoffs, and the concrete implementation details behind RiNET. We'll cover the agent mesh, the LoRA fine-tuning pipeline, the decentralized orchestration layer, and the self-healing infrastructure that keeps the system running.

The Core Architecture

RiNET is not a monolith. It's a distributed mesh of specialized agents, each fine-tuned for a specific cognitive function: reasoning, retrieval, planning, execution, and auditing. Agents communicate via a decentralized message bus (NATS JetStream) and coordinate through a lightweight consensus protocol based on Raft.

┌──────────────────────────────────────────────────┐
│                  RiNET Mesh                       │
│  ┌─────────┐  ┌─────────┐  ┌─────────┐          │
│  │Reasoner │  │Retriever│  │Planner  │ ...       │
│  └────┬────┘  └────┬────┘  └────┬────┘          │
│       │            │            │                │
│       └────────────┼────────────┘                │
│                    │                              │
│        ┌───────────┴───────────┐                  │
│        │   NATS JetStream      │                  │
│        └───────────────────────┘                  │
│                    │                              │
│        ┌───────────┴───────────┐                  │
│        │   Raft Consensus      │                  │
│        └───────────────────────┘                  │
└──────────────────────────────────────────────────┘

Each agent runs as a separate container, with its own LoRA adapter loaded on top of a base model (we use Mistral 7B for most agents, fine-tuned with QLoRA). The base model is shared across agents via a read-only volume mount, while each agent loads its own LoRA weights at startup.

Agent Communication Protocol

Agents communicate via typed messages with a schema:

{
  "id": "uuid",
  "type": "request | response | event",
  "source": "agent-name",
  "target": "agent-name | *",
  "payload": { ... },
  "timestamp": "ISO8601",
  "ttl": 30000
}

Messages are published to NATS subjects like ri.agent.reasoner.request. Agents subscribe to their own subjects and respond. The Raft consensus layer ensures exactly-once delivery and ordering for critical messages (e.g., financial transactions).

Fine-Tuning with LoRA at Scale

RiNET's agents are not generic LLMs. Each agent is fine-tuned on domain-specific data using LoRA adapters. We built a pipeline that:

  1. Collects institutional data (documents, chat logs, databases) via a secure connector.
  2. Tokenizes and filters with a custom deduplication algorithm (MinHash LSH).
  3. Generates training pairs using a teacher model (GPT-4) for instruction tuning.
  4. Trains LoRA adapters with QLoRA on consumer GPUs (RTX 4090 or A6000).

A typical training run:

# Fine-tune a reasoning agent on 10k examples
python train.py \
  --base_model mistralai/Mistral-7B-v0.1 \
  --lora_r 64 \
  --lora_alpha 16 \
  --dataset_path ./data/reasoning_train.jsonl \
  --output_dir ./adapters/reasoner \
  --num_epochs 3 \
  --batch_size 4 \
  --learning_rate 2e-4

Training a single adapter takes ~4 hours on a single RTX 4090. We incrementally update adapters as new data arrives, without retraining the base model.

Adapter Management

We version control LoRA adapters using a custom registry backed by S3-compatible storage. Each adapter has a manifest:

name: reasoner-v2
base_model: mistralai/Mistral-7B-v0.1
lora_r: 64
training_date: 2024-03-15
training_samples: 10240
validation_loss: 0.34
checksum: sha256:abc123...

Agents pull the latest adapter at startup. Rolling updates are handled by the orchestration layer — we can update an adapter without downtime by spinning up a new agent container and switching traffic via NATS.

Orchestration and Self-Healing

RiNET's orchestration layer is built on HashiCorp Nomad, with custom health checks and a decision engine that triggers recovery actions.

Health Checks

Each agent exposes a /health endpoint that returns:

{
  "status": "ok | degraded | down",
  "model_loaded": true,
  "adapter_version": "reasoner-v2",
  "last_inference_ms": 120,
  "error_rate_1m": 0.02
}

Nomad runs these checks every 10 seconds. If an agent fails 3 consecutive checks, the orchestrator:

  1. Marks the agent as unhealthy.
  2. Spins up a replacement container.
  3. Updates NATS subject subscriptions to route traffic to the new instance.
  4. Logs the event to the audit trail.

Decision Engine

For complex failures (e.g., model corruption, data drift), we use a small rule-based engine written in Go:

func evaluateHealth(agent AgentStatus) Action {
    if agent.ErrorRate > 0.1 {
        return Action{RollbackAdapter: agent.AdapterVersion}
    }
    if agent.LastInferenceMs > 5000 {
        return Action{RestartAgent: true}
    }
    if agent.AdapterVersion != expected {
        return Action{UpdateAdapter: expected}
    }
    return Action{Noop: true}
}

This engine runs as a separate Nomad job and can trigger automated rollbacks, restarts, or even full node replacement.

Data Sovereignty

RiNET never sends data outside the institutional boundary. All data — training data, inference inputs, logs — stays on-premises or in a private cloud. We enforce this with:

  • Network policies: Agents can only communicate within the mesh. No outbound internet access except to a private container registry.
  • Encryption at rest and in transit: All data volumes use LUKS encryption. TLS everywhere.
  • Audit trails: Every inference request is logged with a cryptographic hash chain. Tampering is detectable.

Private Model Registry

We run a private Docker registry and a Hugging Face model registry on-premises. Base models are downloaded once and scanned for vulnerabilities. Adapters are built from source and never leave the network.

Real-World Performance

We deployed RiNET in a mid-sized financial institution (500 employees) for 3 months. Here are concrete numbers:

  • Agent count: 12 agents (4 reasoners, 3 retrievers, 2 planners, 2 executors, 1 auditor)
  • Infrastructure: 4 nodes, each with 2x RTX 4090, 64GB RAM, 1TB NVMe
  • Inference latency: p50 = 180ms, p99 = 1.2s (including network)
  • Throughput: 1200 requests/minute peak
  • Uptime: 99.97% (excluding planned maintenance)
  • Self-healing events: 47 automatic recoveries (mostly transient model load failures)
  • Data stored: 2.3 TB (documents, embeddings, logs)

Lesson Learned: Adapter Conflicts

Early on, we discovered that two agents with different adapters sharing the same base model caused memory corruption. The fix: each agent loads the base model into a separate process namespace using unshare. This isolates memory and prevents cross-agent interference.

# Launch agent in isolated namespace
unshare --mount --uts --ipc --pid --fork \
  --mount-proc /usr/local/bin/agent

The Roadmap

RiNET is still evolving. Our next milestones:

  • Federated learning across institutions: Train adapters without sharing raw data.
  • Dynamic agent spawning: Agents that self-replicate based on workload.
  • Formal verification: Prove agent behaviors using TLA+.

We're open-sourcing the core components under the Apache 2.0 license. The repo includes the agent SDK, the orchestration layer, and the adapter management tools.

Final Thoughts

Building a sovereign autonomous OS for institutional cognition is not just about stitching together LLMs. It requires a ground-up rethinking of infrastructure, data governance, and system resilience. RiNET proves it's possible — with the right architecture and a lot of engineering discipline.

If you're building something similar, start with the data sovereignty layer. Everything else depends on it.

#agent-swarms#autonomous-systems#data-sovereignty#institutional-cognition#lora-fine-tuning#self-healing-infrastructure#sovereign-ai
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related