Architecture · technical baseline
A distributed worker system with a durable central coordination layer.
The architecture baseline separates control, placement and durable state from worker-side agent execution. Components and protocols described here are proposed unless explicitly labeled as source-verified.
Control-plane decisions; worker-side execution.
API, workflow, policy and scheduler services coordinate registered VPS workers. Durable state and artifacts support observable, recoverable attempts.
Why a central control plane?
It provides one place to validate requests, apply project policy, rank eligible workers and track task state. The control plane coordinates; it does not turn the platform into a single agent process or live machine monitor.
Why distributed workers?
Agent processes need compute, storage, browser and network resources near the selected execution environment. Workers report capacity and execute bounded attempts while the control plane retains the durable record.
Separate decisions from local runtime duties.
Workers use an outbound connection pattern in the MVP design. A lease identifies the authorized attempt and fencing generation.
Central control plane
Gateway/authentication, project and task management, workflow orchestration, policy, budget reservation, scheduler, VPS registry, state and event interfaces.
Worker runtime
Enrollment identity, resource probes, task claim, sandbox preparation, adapter launch, heartbeat, checkpoint/artifact reporting and idempotent cleanup.
Proposed state backbone
PostgreSQL is the planned source of truth for task state, leases and outbox events. Artifact payloads use separate S3-compatible storage. Redis or NATS should be introduced only if queue/outbox benchmarks justify them.
Transport
HTTPS long-poll plus heartbeat is the initial protocol direction. Workers make outbound TLS connections; public inbound worker ports are not required for this MVP approach.
Each state transition leaves an explicit trace.
Requests become durable tasks, capacity-aware assignments, isolated attempts and checked artifacts.
Recover only when the outcome is safe to recover.
Heartbeat gaps make a worker suspect; lease expiry moves its attempt into reconciliation. A newer fencing generation prevents a late worker from overwriting current state.
Pure or idempotent tasks may be retried. Compatible checkpoints require version and checksum validation. External mutations with uncertain outcomes are held for provider reconciliation or human review—not blindly replayed.
Exactly-once external side effects are not guaranteed unless the external system supplies idempotency or reconciliation support.
Reusable components are not the distributed platform.
ANDIP preserves upstream projects as independent foundations and adds adapter contracts plus newly engineered orchestration infrastructure.
Verified source
Agent Orchestrator source includes local lifecycle, adapter, runtime, worktree and HTTP/event patterns. Jev Ultrafast source includes a browser agent loop, DOM snapshot and guarded actions; its audit reported 31 tests and lint passing.
Verified sourcePlanned ANDIP components
Distributed scheduler, VPS registry, durable lease queue, multi-worker runtime, scoped secrets, centralized policy/budget, artifact contract, recovery and fleet monitoring require additional engineering.
PlannedEngineering assumptions
PostgreSQL lease/outbox design, rootless container execution, S3-compatible artifact storage and HTTPS long-poll worker protocol are architecture choices subject to implementation benchmarks and security validation.
AssumptionFeatures awaiting validation
Cross-VPS capacity placement, durable recovery, isolation, browser profiles, cost metering and the 100-agent workload target need reproducible tests. No ANDIP implementation completion is inferred from upstream source.
Under validationKeep the first system understandable—and measurable.
Why Kubernetes is not required for the initial MVP
The initial design registers already-provisioned Linux VPS and runs a worker service. At this stage, Kubernetes would introduce cluster operations and another control plane without demonstrated need for autoscaling or managed provisioning. Revisit only when test evidence and deployment needs justify it.
Agent supervision is not distributed orchestration
A local harness can create sessions, launch an agent and observe its process. Distributed orchestration must also own worker identity, fleet capacity, leases, queue durability, reservation, policy, artifact persistence and safe recovery across servers.
Architecture baseline source: AGENT_NATIVE_INFRASTRUCTURE_MASTER_WORKFLOW.md, audited 2026-10-08. All operational interfaces remain subject to engineering implementation and change.