At the beginning of our livestream with Microsoft Research engineer Alex Lavaee, he opened an empty project, asked an AI coding agent to build a 3D Subway Surfers-style game, and went back to his presentation.
While Alex explained how to build reliable AI coding workflows, the agent generated characters, 3D objects, movement physics, obstacles and tests. By the end, there was a playable game called Neon Rush.
Watch the completed game at 1:59:18, or explore the source.
But the game was only half the lesson. Alex previously worked on coding agents for Microsoft's Windows team and now researches agentic engineering. His central argument: AI can produce code much faster than humans can inspect it. Proving that code deserves to ship is becoming the hard part.
Correctness is a property of the system around the model, but it's not something the model reports.
This is a practical guide to verification engineering, with linked examples from the livestream, research findings, and prompts you can adapt. The prompts below are Neuron adaptations of the demonstrated practices, not verified verbatim extracts of Alex's deck. Get the original presentation deck and its prompts here.
- Seven rules for trusting coding agents
- Why verification is the new bottleneck
- Testing, verification, validation, and oracles
- Why passing tests isn't enough
- Six principles of verification engineering
- Atomic: executable agent workflows
- What we learned from the game demonstration
- Costs, background work and practical setup
- Security becomes more important as autonomy grows
- A complete workflow you can copy
- What engineers still need to decide
- Everything Alex shared
- Full hyperlinked timecode index: Alex Lavaee's top insights
- Full hyperlinked timecode index: Alex Lavaee's top insights
Seven rules for trusting coding agents
- Define observable success before building.
- Don't let the agent grade its own homework.
- Prefer executable tests to an AI saying something works.
- Test the real application and its important dependencies.
- Save screenshots, logs, recordings, and test results.
- Allocate more verification to riskier changes.
- Keep people responsible for choosing what success means.
Why verification is the new bottleneck
When humans wrote most implementation code, producing the code itself took time. Now multiple agents can build features at once. But when they finish, someone must determine whether they meet the actual requirements.
Alex highlighted OpenAI's harness-engineering report: a small engineering team generated roughly one million lines of code across around 1,500 pull requests. Watch Alex discuss the bottleneck at 15:32.
If an AI writes bugs faster than we can find them, we're also shipping bugs faster.
Testing, verification, validation, and oracles
Testing is building checks for behavior. Verification asks whether the implementation meets agreed criteria. Validation asks whether those criteria solve the user's actual problem.
Imagine a login button that responds when clicked but never creates an authenticated session. The button behavior could pass a narrow test while the product fails its real purpose.
An oracle is whatever judges whether a result passes: a numeric threshold, expected output, reference implementation, automated test, or appropriately qualified reviewer. Watch Alex explain these concepts at 19:27 and our discussion of evals and oracles at 20:43.
Why passing tests isn't enough
Alex compared an agent that changes its tests to a student rewriting the answer key. In a particular METR reward-hacking experiment, telling the model not to cheat left the reward-hacking rate at 80% across the tested runs. That is a specific benchmark result, not a claim that 80% of coding-agent work is cheating. Watch at 24:44.
Alex also discussed ImpossibleBench, where preventing agents from modifying tests blocked a demonstrated cheating route, and UTBoost, which found 345 erroneous patches previously classified as successful by existing SWE-bench tests. Watch at 32:32.
A test can also accidentally codify a bug: if the program incorrectly calculates $100 plus 10% tax as $120, a model that copies current output into the expected result has created a passing test for the bug.
The lesson: derive expectations independently of the implementation and enforce important limits through real permissions and automated gates.
Six principles of verification engineering
Alex introduces the framework at 59:30.
- Specify objective success criteria before the work.
- Separate author and judge.
- Prefer deterministic checks before subjective model judgment.
- Verify in a running environment.
- Save evidence of execution.
- Require human approval for criteria and consequential steps.
1. Define success that you can observe
Alex used a lane-switching example from his Subway Surfers game. Rather than “make movement responsive,” require: With the player in the center lane and a train approaching, pressing left once must move the player into the left lane within 0.25 seconds and the run must continue.
That establishes a starting state, action, observable result, threshold, and check without telling the agent how to implement it.
Adapted prompt: acceptance criteria
Before writing code, define objective acceptance criteria.For each criterion provide:1. Starting state and actor.2. Action.3. Observable expected result.4. Measurable pass/fail threshold.5. Oracle and real environment needed.6. Conditions that require human approval.Do not implement until I approve the criteria.
2. Write a test contract first
An acceptance criterion states what success means. A test contract specifies how to measure it: test level, oracle, real dependencies, threshold, and evidence to preserve. Watch at 1:06:24.
Adapted prompt: test contract
For every approved acceptance criterion:- Select a unit, integration or end-to-end test.- Define the independent expected result.- Specify the real environment and dependencies.- Set measurable pass/fail thresholds.- Name the screenshots, logs or reports to save.If required testing cannot run, explain why and stop.Do not silently swap required real services for mocks.
Atomic's create spec skill can produce architecture, design, testing requirements and unresolved questions. Watch at 1:13:20.
3. Build one vertical slice at a time
A vertical slice is one complete user behavior that you can test from start to finish. For a shopping site: first add an item to cart, then update its quantity, then proceed to checkout. For a bug fix, verify that a new regression test fails against the original buggy version before fixing it. Watch at 1:33:33 and 1:34:14.
Adapted prompt: vertical slices
Implement one approved behavior at a time.Write a focused test, run it before implementation,show the expected failure, implement the behavior,rerun the test and relevant regressions, and save evidence.For bug fixes, prove that the test fails on the original bug.Keep changes small and reviewable.
4. Put deterministic checks at the bottom of the pyramid
Watch Alex's checking pyramid at 1:19:32.
Type checkers, linters and static analysis catch many errors without AI judgment. Unit tests check small components, integration tests check services working together, and end-to-end tests check the actual user's journey. Visual or qualitative judgments can sit at the top, where machines cannot easily enforce a numeric rule.
Alex mentioned tools including Biome, Vitest, Clippy and QLTY. QLTY can flag maintainability issues such as code duplication, nesting and cyclomatic complexity (the number of independent paths through code). Watch at 1:16:00.
Adapted prompt: quality gates
Run all existing checks, then relevant type checks,linting, static analysis, unit, integration and end-to-end tests.Report exact commands, results, failures and skipped checks.Never call a check passed unless it actually ran.
5. Test the real application
A mock is a fake substitute for a dependency. It can be useful in unit testing, but a fake payment-success message doesn't prove that real checkout integration works. Alex recommends using development versions of real services, isolated databases, browsers, and device simulators when the requirement depends on them. Watch at 1:35:17.
Adapted prompt: mock audit
Identify all mocks, fakes, stubs and simulated dependencies.For each, say what it replaces, whether it is appropriatefor the test, and which real integration or end-to-end testis still needed. Propose a safe development test plan.Do not change code yet.
6. Make tests harder to fool
Watch Alex's testing methods overview at 1:38:26.
Property-based testing generates many inputs and checks that required properties always hold (the runner can never move outside three lanes). Metamorphic testing compares related runs (frame rate shouldn't radically change game speed). Differential testing compares new output to a reference. Mutation testing intentionally introduces small bugs to see whether tests detect them. Approved snapshots catch unintended visual changes. Invariants define rules that should always remain true. Seeded replays reproduce the same conditions for debugging.
Alex also suggested each test explain what bug it catches, what change should make it fail, and why existing checks aren't enough. Watch at 1:33:04.
7. Separate the coding agent from the reviewer
A coding agent may have spent an hour justifying its decisions in the same conversation. Asking it to review that work can carry those assumptions forward.
Alex recommends an independent reviewer with a fresh context window and access to the specification, diff, test results, and saved evidence. A fresh context doesn't make AI infallible, but can reduce one source of review bias. Watch at 1:41:02.
Adapted prompt: independent review
Review another agent's work in a fresh context.Use the approved specification as source of truth.Inspect the diff, tests and saved evidence.Look for unmet criteria, weak or manipulated tests,regressions, security issues and maintainability problems.Give specific findings with evidence and severity.Do not change the code during this review.
8. Preserve evidence of completion
When an agent finishes, collect the test output, screenshots, browser recordings, logs, metrics and relevant commit details. This lets a human inspect evidence of behavior rather than simply trusting a completion message. Watch at 1:01:00 and 1:56:17.
9. Scale the verification budget with risk
A spacing change, a new reporting feature and a payments overhaul need different amounts of scrutiny. Alex described allocating 30 minutes, an hour, or even five hours to verifying a substantial change. Watch at 1:06:54. The useful measure is not merely how many tests run, but whether the evidence matches the consequence of failure.
Atomic: executable agent workflows
Atomic is Alex's open-source runtime for agentic engineering. It can turn natural-language process instructions into TypeScript workflows containing workers, checks, parallel operations, repair loops, approval gates and saved artifacts. Watch its workflow graph at 35:09.
A state machine moves through defined stages such as Plan -> Build -> Test -> Review -> Complete. If tests fail, it can return to Build; if approval is required, it waits. Recursive workflows can contain smaller workflows. Atomic also includes an Intercom mechanism for communication among participating agents. Watch the architecture at 35:56 and communication at 38:51.
Alex proposed a progression from prompts -> skills -> workflows: prompts request work, skills package reusable knowledge, and workflows make repeatable process logic executable. Watch at 1:24:06.
Adapted prompt: turn a skill into a workflow
Review this repeated process and identify:inputs, sequential and parallel stages, tools,checks, retry paths, approval gates, outputs,saved evidence, and stop conditions.Present the architecture before implementation.Where supported, implement the repeatable controllogic with the workflow SDK instead of prose alone.
What we learned from the game demonstration
Alex launched the project around 9:35. Atomic handled visual assets, conversion into 3D meshes, gameplay mechanics and verification. A mesh is the geometric structure representing a 3D object.
The agent created performance and game-state measurements, and Alex demonstrated bounding-box checks for locating objects. A bounding box identifies an object's approximate spatial region for visual checks and collision logic. See early gameplay at 26:32 and bounding boxes at 1:27:25.
He also discussed multimodal checks: ask an image-capable model to inspect a screenshot and determine whether the runner is in the left lane. These are useful where visual interpretation matters, but they remain model judgments rather than guarantees. Watch at 1:55:23.
The important demonstration was not simply that an agent could generate a game. It could also construct mechanisms for measuring the game's behavior.
Costs, background work and practical setup
Running extra models and reviewers can consume more tokens, but Alex argued that total cost should include failed attempts and human review time. The livestream didn't establish controlled cost comparisons, so treat these as an engineering hypothesis rather than proven savings. Watch the cost discussion at 41:10.
For long-running work, Alex discussed terminal session tools, durable workflows, separate Git worktrees (isolated working directories linked to one repository), and using GitHub issues as a task queue. The agent can complete an issue and prepare a pull request for later review. Watch the GitHub workflow at 57:51.
Atomic supports multiple model providers and authentication routes; check its documentation for current availability. Alex said it works for both greenfield (new) and brownfield (existing) projects. Watch at 52:19.
Alex said he had not submitted Atomic to Terminal-Bench. He described favorable engineering anecdotes, but those are not standardized independent benchmarks. Watch at 49:54.
Security becomes more important as autonomy grows
An agent that can run shell commands, use cloud credentials and modify a repository can also do damage. Use isolated development environments, least-privilege credentials, development accounts, spending limits, logs and explicit approval for destructive actions. Alex discussed credential and cost safeguards at 1:49:15.
Atomic's repository documentation warns that its runtime itself is not a built-in security sandbox. Plan the surrounding environment accordingly.
A complete workflow you can copy
This is a Neuron-authored synthesis, not a verbatim deck prompt:
TASK: [Describe the feature or fix.]
Before implementation:1. Inspect the relevant code and environment.2. Define objective, observable acceptance criteria.3. Identify what must not change.4. Build a test contract with independent oracles.5. Identify required real environments and dependencies.6. Request my approval of the specification.
After approval:7. Implement one vertical slice at a time.8. Prove new bug-regression tests fail on the old bug.9. Run type checks, linting, tests and static analysis.10. Exercise actual integrations in safe test environments.11. Save browser recordings, screenshots, logs and reports.12. Ask a reviewer with fresh context to inspect the work.13. Fix review findings and rerun verification.14. List what passed, failed or could not be checked.15. Request human approval for consequential operations.
Never weaken protected tests or silently changeacceptance criteria merely to get a passing result.Do not claim checks passed when they were not executed.
What engineers still need to decide
Verification can prove compliance with a set of checks, but validation is still needed to decide whether the checks represent the right user problem. Better models make more sophisticated implementations possible. Better tooling makes their behavior more observable and auditable. Neither improvement eliminates human responsibility for defining success and acceptable risk.
Watch Alex on agentic engineering and vibe coding at 1:39:32.
By the final demo at 1:59:18, the agent had built a playable game. The question for the next wave of software isn't merely how fast an agent can write code. It's whether verification can become cheap and reliable enough to keep pace.
Everything Alex shared
- Full livestream with Corey and Grant
- Alex's presentation deck, including original prompts
- Subway Surfers / Neon Rush source code
- Atomic's open-source repository
- Agentic Engineering Masterclass 1
- Agentic Engineering Masterclass 2
- Mixture of Experts with Alex and Norin Lavaee
Related research: OpenAI Harness Engineering, METR reward hacking, ImpossibleBench, and UTBoost.
Full hyperlinked timecode index: Alex Lavaee's top insights
If you want to jump straight to the most useful moments from the full conversation, this is the complete index of the major ideas, practical instructions, research references, and memorable demonstrations we pulled from the transcript.
The core thesis: writing code is getting cheap, checking it isn't
- (4:01) Alex explains that conventional coding agents struggle as projects and codebases get larger; he connects this to his earlier work on internal agents for Microsoft Windows.
- (5:19) The central idea: when an agent says it is done, that is a claim. Alex says correctness is a property of the system around the model, not something the model can simply report about itself.
- (5:59) Alex breaks verification engineering into three phases: define criteria before work, guide implementation with independent checks, then run the finished system and save evidence.
- (8:18) Alex introduces Atomic as an open-source runtime for verifiable coding-agent workflows rather than only a chat-style coding interface.
- (9:35) The live experiment begins: Atomic starts building a Subway Surfers-style game from an empty project while the interview continues.
- (10:38) Atomic uses a prompt-engineering skill to turn a rough request into clearer instructions for the model.
- (11:14) Alex explains the different AI capabilities involved in the game: image generation, 2D-to-3D conversion, coding, physics, and testing.
- (13:00) Atomic can use one model as the orchestrator while routing other stages to different models based on the job and cost.
- (13:46) The game project illustrates reusable workflow stages for assets, movement, collisions, testing, and verification.
- (15:32) Alex argues that human QA is becoming the bottleneck as coding agents increase software throughput.
- (16:24) Corey compares the problem with AI-generated mathematics, where models can generate results faster than people can carefully review them.
- (19:01) Alex connects this to loop engineering and graph engineering: reliable background workflows are needed if people are going to supervise many agents at once.
Tests, oracles, and why passing is not proof
- (19:27) Alex separates test engineering, verification, and validation. A system can satisfy a test while still failing the user's real need.
- (20:43) Alex defines an oracle as whatever decides whether a result is correct; expected values, references, properties, humans, and evals can all serve as oracles.
- (21:15) High code coverage can still hide weak tests. An agent can generate many tests without checking the behaviors that actually matter.
- (22:19) Alex frames verification as a feedback loop: the system acts, measures the result, and feeds that measurement into the next action.
- (23:51) Evaluators can return richer feedback than a pass/fail score; Alex discusses explanations, multiple verifiers, and optimization loops.
- (24:44) Alex discusses METR's reward-hacking experiment, where adding “please do not cheat” did not reduce the measured reward-hacking rate in that task.
- (25:26) Alex cites ImpossibleBench examples where agents modified tests or otherwise satisfied the grader instead of the intended specification.
- (25:59) Read-only permissions stopped one demonstrated cheating path. The takeaway: system-level controls are stronger than asking nicely in a prompt.
- (26:32) The game begins producing measurable runtime data and physics checks, including position and contact information.
- (28:15) Alex explains empirical verification: run the actual application, perform actions, and observe what happens.
- (29:24) A separate language model can act as a reviewer when deterministic checks cannot settle the question.
- (31:27) Atomic can use independent reviewers in separate contexts instead of relying only on the same agent that wrote the code.
- (32:32) Alex cites UTBoost finding 345 erroneous SWE-bench patches that had previously passed existing tests.
- (33:32) AI-generated tests can learn the current buggy behavior as the expected answer, turning a bug into a passing test.
How Atomic structures autonomous work
- (35:09) Alex shows Atomic's recursive state-machine workflow graph, where different nodes represent separate stages, tools, or agent harnesses.
- (35:56) Alex explains recursive language-model ideas: a model can use programs and additional model calls to manage work too large for one context window.
- (37:18) Atomic workflows can preserve state, allowing interrupted runs to resume instead of starting from scratch.
- (38:14) Users can inspect workflow stages, steer a running workflow, and answer questions during execution.
- (38:51) Atomic supports parallel workflows, periodic orchestrator “heartbeats,” and Intercom communication among agents.
- (40:19) Alex argues that peer-to-peer communication can support richer multi-agent collaboration than supervisor-only structures.
- (41:10) Alex argues that total cost should include human steering and failed loops, not only token spend. The livestream does not provide a controlled benchmark for this claim.
- (43:18) Alex says he uses Atomic to build and test Atomic itself, including browser automation, simulators, terminals, and computer use.
- (44:07) Alex demonstrates Herder, an agent-oriented terminal-session tool for persistent background work and status visibility.
- (45:34) Better models do not remove the value of better runtimes and tools; Alex argues the two compound each other.
- (47:39) Narrowly scoped workflow stages can make behavior easier to control and verify than one giant general-purpose agent context.
- (48:49) Alex shows Atomic's login flow and explains support for multiple model providers and authentication methods.
- (49:54) Alex says he prioritizes real-world engineering case studies over heavily optimized coding benchmarks such as Terminal-Bench.
- (50:56) Alex shares anecdotal reports from engineers using Atomic at large companies, including fewer regressions and less manual review; these are user reports, not controlled benchmark results.
- (52:19) Atomic is designed for both greenfield projects and large existing codebases.
- (53:05) Alex recommends explicit workflow stages that launch the actual browser, app, simulator, or device when native behavior needs verification.
How to manage long-running agents
- (54:36) The game is already playable enough to show running animation and lane changes, but important pieces are still incomplete.
- (55:34) Alex predicts agents will increasingly work for hours or days, making intermediate evidence and checkpoints more important.
- (56:25) Alex discusses routing agent questions and status updates through tools such as Slack or Microsoft Teams via custom integrations.
- (57:51) Alex uses GitHub issues as a task queue and pull requests as the review surface for background agent work.
- (58:50) Atomic's prompt-engineering workflow encourages explicit success criteria, forbidden changes, and stop conditions.
The six verification principles
- (59:30) Principle one: define objective success criteria before the implementation starts.
- (1:00:00) Principle two: the author should not be the sole judge; prefer independent reviewers and deterministic checks.
- (1:01:00) Principles three through five: verify the running system, prefer executable checks where possible, and save evidence such as logs, screenshots, and recordings.
- (1:02:17) Principle six: humans approve important criteria and consequential steps; Alex also introduces verification-time scaling.
How to define “done”
- (1:03:58) Alex gives the lane-switching acceptance criterion: start in the middle lane, press left once, reach the left lane within 0.25 seconds, and continue running.
- (1:04:59) Acceptance criteria can be visual. The requirement can show the end behavior without prescribing the implementation.
- (1:05:55) Alex recommends generating and reviewing acceptance criteria before code changes begin.
- (1:06:24) Criteria define what success means; the test contract defines how each criterion will be checked.
- (1:06:54) Atomic can allocate 30 minutes, an hour, or several hours to verification and manage that time budget during the workflow.
- (1:08:32) Alex says persistent workflow history can improve estimates for how long familiar tasks may take.
- (1:09:28) Alex's test-contract checklist: every criterion needs an oracle, expected values should come from the specification, real dependencies should be identified, and thresholds should be measurable.
- (1:10:02) If a real dependency cannot be tested, stop and report the blocker instead of silently replacing it with a mock.
- (1:13:20) Atomic's generated specifications can include architecture diagrams, implementation design, testing requirements, and open questions.
- (1:14:59) Alex recommends enforcing important rules with executable tools instead of relying solely on AGENTS.md or CLAUDE.md instructions.
- (1:16:00) Use the repository's established checks first; Alex mentions type checkers, Biome, Vitest, Clippy, and QLTY.
- (1:17:25) Passing today's tests is different from producing maintainable code; Alex recommends checking duplication, complexity, nesting, and other long-term quality signals.
- (1:17:54) Alex recommends Agent Browser for browser automation and other computer-use tooling for non-browser environments.
- (1:19:32) The pyramid of deterministic checking: types and linting, unit tests, integration tests, end-to-end tests, then model judgment for fuzzy questions.
From prompts to skills to workflows
- (1:24:06) Alex proposes the progression “prompts -> skills -> workflows.”
- (1:24:36) If a skill repeatedly describes steps, branches, retries, and approvals, Alex argues that the repeatable control logic should increasingly become executable.
- (1:25:13) Atomic workflows are TypeScript programs built with its workflow SDK.
- (1:26:04) Workflows can define inputs, outputs, tasks, parallel stages, tools, and separate execution contexts.
- (1:27:25) The game now contains object-detection-style bounding boxes that help the system reason about objects and collisions.
Stronger testing and independent review
- (1:33:04) Every test should justify its existence: what behavior does it protect, which change would make it fail, and why are existing tests insufficient?
- (1:33:33) For a bug fix, prove the new regression test fails on the original buggy implementation before trusting it.
- (1:34:14) Alex recommends vertical test-driven-development slices: one behavior, one failing test, one implementation step, then repeat.
- (1:35:17) Alex strongly favors testing real services for behaviors that depend on them, rather than treating mocks as sufficient proof.
- (1:36:38) Audit the test suite for mocks and identify where real integration or end-to-end testing is still needed.
- (1:37:18) Workflow stages should only advance after relevant verification gates pass.
- (1:38:26) Alex explains property-based tests, metamorphic tests, differential tests, and mutation testing as stronger oracle strategies.
- (1:38:56) Additional strategies include approved snapshots, invariants, contracts, and seeded replays.
- (1:39:32) Alex distinguishes agentic engineering from vibe coding by the engineer's control over specification, verification, and review.
- (1:41:02) A reviewing model should ideally begin with a fresh context rather than inheriting the worker's whole conversational history.
- (1:41:54) Fresh-context review can still be implemented with subagents if you deliberately pass only the information needed for review.
- (1:42:09) Independent contexts may cost more tokens because prior context cannot always be reused, but Alex considers that tradeoff worthwhile for stronger review.
- (1:44:13) Performance and maintainability need explicit verification; technically functional software can still deliver a poor experience.
- (1:45:48) Alex cautions that autonomous coding is still early; impressive demonstrations do not prove reliability across every engineering problem.
- (1:46:17) Verifying the wrong thing is still inadequate verification. Meaningful criteria matter more than simply executing a check.
Security, credentials, and multimodal verification
- (1:48:16) Alex recommends development, dogfood, and canary environments so changes can be exercised before broad production rollout.
- (1:49:15) Autonomous workflows need deliberate credential management and safe access boundaries.
- (1:50:25) Set budget limits on services and credentials an agent can use; autonomous systems can generate expensive volumes of requests quickly.
- (1:51:40) Alex recommends environment files, password managers, and cloud key vaults rather than exposing secrets directly in prompts or code.
- (1:53:57) Fast multimodal decision models can serve as narrow verifiers for visual conditions.
- (1:55:23) In the game demo, a model checks a screenshot and estimates whether the runner is in the left lane.
- (1:56:00) Alex describes model-driven fuzz testing, where an agent explores different game interactions rather than only following a fixed happy path.
- (1:56:17) A verification record should tie the candidate commit, execution environment, checks, and saved evidence together.
The final game and what comes next
- (1:59:18) Alex demonstrates the completed Neon Rush game with three-lane movement, trains, obstacles, animation, music, and visual assets.
- (1:59:38) The game uses procedurally generated object locations rather than a single fixed level.
- (2:00:54) Alex expects open-source models to become increasingly viable for agentic engineering, while local inference remains hardware-intensive.
- (2:02:36) Alex discusses high-memory Apple Silicon as one option for local model inference.
- (2:03:05) Multiple large local models and long contexts create practical memory and throughput bottlenecks.
- (2:03:32) Advertised context-window size does not guarantee that running at that size locally will be practical.
- (2:05:00) Alex lists Atomic priorities: reliability, real production use cases, community feedback, and more contributed workflows.
- (2:05:30) Grant suggests a graphical interface for Atomic to make its capabilities more approachable to non-terminal users.
- (2:05:50) Alex says Atomic's runtime can support custom frontends built through its SDK.
- (2:06:19) Alex discusses integrations that could let developers invoke Atomic from existing coding environments instead of switching interfaces.
- (2:07:35) Alex closes by inviting feedback and pointing viewers toward Atomic's community and Mixture of Experts.
Full hyperlinked timecode index: Alex Lavaee's top insights
If you want to jump straight to the most useful moments from the full conversation, this is the complete index of the major ideas, practical instructions, research references, and memorable demonstrations we pulled from the transcript.
The core thesis: writing code is getting cheap, checking it isn't
- (4:01) Alex explains that conventional coding agents struggle as projects and codebases get larger; he connects this to his earlier work on internal agents for Microsoft Windows.
- (5:19) The central idea: when an agent says it is done, that is a claim. Alex says correctness is a property of the system around the model, not something the model can simply report about itself.
- (5:59) Alex breaks verification engineering into three phases: define criteria before work, guide implementation with independent checks, then run the finished system and save evidence.
- (8:18) Alex introduces Atomic as an open-source runtime for verifiable coding-agent workflows rather than only a chat-style coding interface.
- (9:35) The live experiment begins: Atomic starts building a Subway Surfers-style game from an empty project while the interview continues.
- (10:38) Atomic uses a prompt-engineering skill to turn a rough request into clearer instructions for the model.
- (11:14) Alex explains the different AI capabilities involved in the game: image generation, 2D-to-3D conversion, coding, physics, and testing.
- (13:00) Atomic can use one model as the orchestrator while routing other stages to different models based on the job and cost.
- (13:46) The game project illustrates reusable workflow stages for assets, movement, collisions, testing, and verification.
- (15:32) Alex argues that human QA is becoming the bottleneck as coding agents increase software throughput.
- (16:24) Corey compares the problem with AI-generated mathematics, where models can generate results faster than people can carefully review them.
- (19:01) Alex connects this to loop engineering and graph engineering: reliable background workflows are needed if people are going to supervise many agents at once.
Tests, oracles, and why passing is not proof
- (19:27) Alex separates test engineering, verification, and validation. A system can satisfy a test while still failing the user's real need.
- (20:43) Alex defines an oracle as whatever decides whether a result is correct; expected values, references, properties, humans, and evals can all serve as oracles.
- (21:15) High code coverage can still hide weak tests. An agent can generate many tests without checking the behaviors that actually matter.
- (22:19) Alex frames verification as a feedback loop: the system acts, measures the result, and feeds that measurement into the next action.
- (23:51) Evaluators can return richer feedback than a pass/fail score; Alex discusses explanations, multiple verifiers, and optimization loops.
- (24:44) Alex discusses METR's reward-hacking experiment, where adding “please do not cheat” did not reduce the measured reward-hacking rate in that task.
- (25:26) Alex cites ImpossibleBench examples where agents modified tests or otherwise satisfied the grader instead of the intended specification.
- (25:59) Read-only permissions stopped one demonstrated cheating path. The takeaway: system-level controls are stronger than asking nicely in a prompt.
- (26:32) The game begins producing measurable runtime data and physics checks, including position and contact information.
- (28:15) Alex explains empirical verification: run the actual application, perform actions, and observe what happens.
- (29:24) A separate language model can act as a reviewer when deterministic checks cannot settle the question.
- (31:27) Atomic can use independent reviewers in separate contexts instead of relying only on the same agent that wrote the code.
- (32:32) Alex cites UTBoost finding 345 erroneous SWE-bench patches that had previously passed existing tests.
- (33:32) AI-generated tests can learn the current buggy behavior as the expected answer, turning a bug into a passing test.
How Atomic structures autonomous work
- (35:09) Alex shows Atomic's recursive state-machine workflow graph, where different nodes represent separate stages, tools, or agent harnesses.
- (35:56) Alex explains recursive language-model ideas: a model can use programs and additional model calls to manage work too large for one context window.
- (37:18) Atomic workflows can preserve state, allowing interrupted runs to resume instead of starting from scratch.
- (38:14) Users can inspect workflow stages, steer a running workflow, and answer questions during execution.
- (38:51) Atomic supports parallel workflows, periodic orchestrator heartbeats, and Intercom communication among agents.
- (40:19) Alex argues that peer-to-peer communication can support richer multi-agent collaboration than supervisor-only structures.
- (41:10) Alex argues that total cost should include human steering and failed loops, not only token spend. The livestream does not provide a controlled benchmark for this claim.
- (43:18) Alex says he uses Atomic to build and test Atomic itself, including browser automation, simulators, terminals, and computer use.
- (44:07) Alex demonstrates Herder, an agent-oriented terminal-session tool for persistent background work and status visibility.
- (45:34) Better models do not remove the value of better runtimes and tools; Alex argues the two compound each other.
- (47:39) Narrowly scoped workflow stages can make behavior easier to control and verify than one giant general-purpose agent context.
- (48:49) Alex shows Atomic's login flow and explains support for multiple model providers and authentication methods.
- (49:54) Alex says he prioritizes real-world engineering case studies over heavily optimized coding benchmarks such as Terminal-Bench.
- (50:56) Alex shares anecdotal reports from engineers using Atomic at large companies, including fewer regressions and less manual review; these are user reports, not controlled benchmark results.
- (52:19) Atomic is designed for both greenfield projects and large existing codebases.
- (53:05) Alex recommends explicit workflow stages that launch the actual browser, app, simulator, or device when native behavior needs verification.
How to manage long-running agents
- (54:36) The game is already playable enough to show running animation and lane changes, but important pieces are still incomplete.
- (55:34) Alex predicts agents will increasingly work for hours or days, making intermediate evidence and checkpoints more important.
- (56:25) Alex discusses routing agent questions and status updates through tools such as Slack or Microsoft Teams via custom integrations.
- (57:51) Alex uses GitHub issues as a task queue and pull requests as the review surface for background agent work.
- (58:50) Atomic's prompt-engineering workflow encourages explicit success criteria, forbidden changes, and stop conditions.
The six verification principles
- (59:30) Principle one: define objective success criteria before the implementation starts.
- (1:00:00) Principle two: the author should not be the sole judge; prefer independent reviewers and deterministic checks.
- (1:01:00) Principles three through five: verify the running system, prefer executable checks where possible, and save evidence such as logs, screenshots, and recordings.
- (1:02:17) Principle six: humans approve important criteria and consequential steps; Alex also introduces verification-time scaling.
How to define “done”
- (1:03:58) Alex gives the lane-switching acceptance criterion: start in the middle lane, press left once, reach the left lane within 0.25 seconds, and continue running.
- (1:04:59) Acceptance criteria can be visual. The requirement can show the end behavior without prescribing the implementation.
- (1:05:55) Alex recommends generating and reviewing acceptance criteria before code changes begin.
- (1:06:24) Criteria define what success means; the test contract defines how each criterion will be checked.
- (1:06:54) Atomic can allocate 30 minutes, an hour, or several hours to verification and manage that time budget during the workflow.
- (1:08:32) Alex says persistent workflow history can improve estimates for how long familiar tasks may take.
- (1:09:28) Alex's test-contract checklist: every criterion needs an oracle, expected values should come from the specification, real dependencies should be identified, and thresholds should be measurable.
- (1:10:02) If a real dependency cannot be tested, stop and report the blocker instead of silently replacing it with a mock.
- (1:13:20) Atomic's generated specifications can include architecture diagrams, implementation design, testing requirements, and open questions.
- (1:14:59) Alex recommends enforcing important rules with executable tools instead of relying solely on AGENTS.md or CLAUDE.md instructions.
- (1:16:00) Use the repository's established checks first; Alex mentions type checkers, Biome, Vitest, Clippy, and QLTY.
- (1:17:25) Passing today's tests is different from producing maintainable code; Alex recommends checking duplication, complexity, nesting, and other long-term quality signals.
- (1:17:54) Alex recommends Agent Browser for browser automation and other computer-use tooling for non-browser environments.
- (1:19:32) The pyramid of deterministic checking: types and linting, unit tests, integration tests, end-to-end tests, then model judgment for fuzzy questions.
From prompts to skills to workflows
- (1:24:06) Alex proposes the progression “prompts → skills → workflows.”
- (1:24:36) If a skill repeatedly describes steps, branches, retries, and approvals, Alex argues that the repeatable control logic should increasingly become executable.
- (1:25:13) Atomic workflows are TypeScript programs built with its workflow SDK.
- (1:26:04) Workflows can define inputs, outputs, tasks, parallel stages, tools, and separate execution contexts.
- (1:27:25) The game now contains object-detection-style bounding boxes that help the system reason about objects and collisions.
Stronger testing and independent review
- (1:33:04) Every test should justify its existence: what behavior does it protect, which change would make it fail, and why are existing tests insufficient?
- (1:33:33) For a bug fix, prove the new regression test fails on the original buggy implementation before trusting it.
- (1:34:14) Alex recommends vertical test-driven-development slices: one behavior, one failing test, one implementation step, then repeat.
- (1:35:17) Alex strongly favors testing real services for behaviors that depend on them, rather than treating mocks as sufficient proof.
- (1:36:38) Audit the test suite for mocks and identify where real integration or end-to-end testing is still needed.
- (1:37:18) Workflow stages should only advance after relevant verification gates pass.
- (1:38:26) Alex explains property-based tests, metamorphic tests, differential tests, and mutation testing as stronger oracle strategies.
- (1:38:56) Additional strategies include approved snapshots, invariants, contracts, and seeded replays.
- (1:39:32) Alex distinguishes agentic engineering from vibe coding by the engineer's control over specification, verification, and review.
- (1:41:02) A reviewing model should ideally begin with a fresh context rather than inheriting the worker's whole conversational history.
- (1:41:54) Fresh-context review can still be implemented with subagents if you deliberately pass only the information needed for review.
- (1:42:09) Independent contexts may cost more tokens because prior context cannot always be reused, but Alex considers that tradeoff worthwhile for stronger review.
- (1:44:13) Performance and maintainability need explicit verification; technically functional software can still deliver a poor experience.
- (1:45:48) Alex cautions that autonomous coding is still early; impressive demonstrations do not prove reliability across every engineering problem.
- (1:46:17) Verifying the wrong thing is still inadequate verification. Meaningful criteria matter more than simply executing a check.
Security, credentials, and multimodal verification
- (1:48:16) Alex recommends development, dogfood, and canary environments so changes can be exercised before broad production rollout.
- (1:49:15) Autonomous workflows need deliberate credential management and safe access boundaries.
- (1:50:25) Set budget limits on services and credentials an agent can use; autonomous systems can generate expensive volumes of requests quickly.
- (1:51:40) Alex recommends environment files, password managers, and cloud key vaults rather than exposing secrets directly in prompts or code.
- (1:53:57) Fast multimodal decision models can serve as narrow verifiers for visual conditions.
- (1:55:23) In the game demo, a model checks a screenshot and estimates whether the runner is in the left lane.
- (1:56:00) Alex describes model-driven fuzz testing, where an agent explores different game interactions rather than only following a fixed happy path.
- (1:56:17) A verification record should tie the candidate commit, execution environment, checks, and saved evidence together.
The final game and what comes next
- (1:59:18) Alex demonstrates the completed Neon Rush game with three-lane movement, trains, obstacles, animation, music, and visual assets.
- (1:59:38) The game uses procedurally generated object locations rather than a single fixed level.
- (2:00:54) Alex expects open-source models to become increasingly viable for agentic engineering, while local inference remains hardware-intensive.
- (2:02:36) Alex discusses high-memory Apple Silicon as one option for local model inference.
- (2:03:05) Multiple large local models and long contexts create practical memory and throughput bottlenecks.
- (2:03:32) Advertised context-window size does not guarantee that running at that size locally will be practical.
- (2:05:00) Alex lists Atomic priorities: reliability, real production use cases, community feedback, and more contributed workflows.
- (2:05:30) Grant suggests a graphical interface for Atomic to make its capabilities more approachable to non-terminal users.
- (2:05:50) Alex says Atomic's runtime can support custom frontends built through its SDK.
- (2:06:19) Alex discusses integrations that could let developers invoke Atomic from existing coding environments instead of switching interfaces.
- (2:07:35) Alex closes by inviting feedback and pointing viewers toward Atomic's community and Mixture of Experts.
Thanks for scrolling this far down! You're a real one. :D