October 2026
TraderUK - Multi Ai Agent
a doer agent writes the code, a PR-checker reviews it against QA gates
Agentic AI
Market Intelligence
Role
AI Engineer
Timeline
October 2026
team
platform
Cursor CLI, Claude, OpenAI, Exa, Harness Engineering

TraderUK watches markets, lets typed AI agents research what it finds, and passes anything that might become a trade through deterministic code that no model can override. Agents return judgments. Code decides eligibility, price levels, stops, targets, size and risk, and a model can't send an order or invent a risk number.
a good chunk of the work is plumbing that stops the system claiming more than the evidence supports.
Constraints
[MISSING: the limits you worked within.]
How It's Built
the codebase is Python 3.11, split into layers: data, an experiment engine, agentic research, policy/risk/portfolio, execution, observability and a control plane. underneath sit eight SQL migrations. they default the kill switch to on, make the event tables append-only with triggers, and define the database roles and grants.
on the research side, Hermes runs as isolated research runners. they claim jobs through a typed contract, work from a frozen snapshot and hand back a result. jobs are persistent with atomic leases, accepted results are immutable, and if an agent times out the opportunity passes through tagged as pass-through. it never counts as a successful analysis. requests between services are scoped, signed (bearer plus HMAC) and replay-protected. Hermes gets no broker credentials and no execution tool, and by design the execution path keeps working if Hermes or a model goes down.
the evaluation setup compares four arms (A, B1, B2, C) on identical baseline opportunities, with B1 as the primary challenger.
Technical Bits
FastAPI, Pydantic, psycopg 3 and httpx on PostgreSQL and Redis, hosted on Railway. the infrastructure is defined in TypeScript, and creating a live environment is blocked in the code itself. the Docker image runs as a non-root user from a non-editable install. Ruff and strict mypy are on, and GitHub Actions CI on Python 3.11 runs ops validation, import-boundary checks, a secret scan, lint, types and tests. at the audited revision (4 October 2026) that was 258 tests passing and 18 skipped. the skips are the real Postgres role tests, live Jev and a real OANDA practice test, so those gates are still unproven.
Where It Honestly Stands
checked: the Phase 1 reporting semantics are implemented and covered by tests. on the 4 October 2026 audit, staging had the API, research API, worker, PostgreSQL and Redis deployed. signed ops reads were accepted, and unsigned reads, replays and bad signatures were rejected. the kill switch was on, live auto was off, and /ready returned 503 saying not ready, which was the correct answer.
not there yet: there's no real US-equity data pipeline. Massive, Benzinga and SEC EDGAR are the proposed sources and none is wired in. the Jev evaluation code exists, but its bundled 200-label set is a synthetic fixture and human adjudication is pending. the staging execution worker isn't deployed. the audit also found the deployed database URLs still using the owner login instead of separate least-privilege roles, and a dashboard that was publicly reachable and showed some hard-coded values. those are the next fixes.
no profitability claim, no accuracy number, no calibration claim. there's no forward evidence for any of them yet, and the platform is built to say so.
