Operating lead of this lab. I run a multi‑machine agent fleet in production and publish how it is governed.
Most teams can get one AI agent working. Very few can run a fleet of them safely. The hard part is not the model: it is deciding what an agent may do on its own, making agents check each other, and making their failures visible instead of silent. That is the layer I build, and I publish it, because you should be able to check it before you trust it.
→ Nine 2026 preprints, each with a DOI and the code
→ Numbers you can check ·
Record ·
Published work ·
Talks ·
Contact
Every figure below links to the thing that produces it. Counts were re‑measured on the dates given. Nothing is rounded up.
38 code and documentation fixes plus 12 curated-list entries, among them UK AISI
inspect_ai, Qwen Code, the MCP Go SDK, Pydantic Logfire and the Gemini cookbook.
Claude Code, Codex, Grok, Gemini, GLM and Mistral, on a shared skill shelf with an incident journal. The rulebook the fleet actually runs on is published, not described.
57 of them were the lab's recurring online Collective Call, and most of those drew single digits: 433 registrations across all 58, of which 245 came from the one event held in person, AI Agents Pad: VCs & Investors Private VIP, Paris, 9 April 2025. So the number is a record of showing up every week for a year, not of filling rooms. The in-person side is where this lab is thin, and the calendar is public, so you can check that yourself.
Journal articles, conference papers, book chapters and preprints. The full record is public and includes one entry listed as retracted, because the publisher withdrew the whole proceedings volume it appeared in. It is listed there rather than quietly dropped.
An independent lab working on multi-agent coordination, agent governance and how agent failures are made visible. I run the fleet the research comes out of, and write its operating manual in public.
A hired operating executive, not a shareholder: I ran delivery for a software house and a startup incubator, and scaled a distributed engineering organisation to 40+ developers across APAC. Built Solidity curricula, ran hackathons, cohorts and incubations. Advised enterprise subsidiaries of Foxconn and ANA Airlines.
Credit and risk systems for a cross-border lending product.
Five years building CRM and ERP on Microsoft Navision and Salesforce, after running hardware sales for the region. The operating toolkit I automate with agents today started here.
Programs in Python and C++. MSc in computer and information systems security, National Research Nuclear University MEPhI; PhD in Education (Information Technologies), 2021. Working languages: English, Russian, Ukrainian, Polish.
Nine preprints on Zenodo, deposited between 7 September and 6 October 2026, CC BY 4.0, each with a permanent DOI and a full-text PDF, seven of them with the repository the paper describes. These are experience reports from a system that actually runs, not proposals.
All nine with abstracts and PDFs · peer-reviewed record on the academic profile
I speak about what breaks in a running agent fleet, with the logs on screen. One night my agents filed 126 incident reports about a network that was fine; the power log showed the Mac had fallen asleep 69 times. That is the kind of material: the failure, the log line that settled it, and the rule we run now because of it.
Speaker page: talk abstracts, formats and how to invite me
The lab's artifacts are open and reproducible. The fastest way to judge the work is to clone one and try to break it.
Full artifact list on the lab page
Based in Palo Alto, California. Work is distributed across the Bay Area, Lisbon and APAC. I hold Polish (EU) citizenship and an active US O‑1; no employer sponsorship is required in either market.
Book 30 minutes · [email protected] · Telegram @tonydzi · WhatsApp +1 341 222 9178 · X @Tony_Stef_ · GitHub
If you are weighing whether to put agents in charge of something that matters, the useful first conversation is about which decisions you are willing to let them make alone. Bring that question and we will get somewhere in half an hour.