orch: ticket-based work with coding agents
orch is a way of working with coding agents: agents do the work through Markdown tickets, while a person approves the requirements, the plan and the result, answers the critical questions, and decides based on evidence rather than the agent's summary.
Code on GitHub: github.com/severinlindenmann/orch-core · open source, Apache 2.0
- Tickets are the order and the memory. They live as files in the repository, next to the code.
- Three gates keep every decision human: requirements, plan, verdict. An approval binds the exact text that was shown.
- Agents prove their work. A ticket reaches testing only with acceptance criteria, a ticked task list and linked evidence.
This page is about the concept. The implementation is open source as orch-core (Apache 2.0); its tagline sums it up: Let coding agents do the work. Keep every decision yours.
Why do coding agents need a workflow with hard edges?
Agents are fast and confident. Left alone, they will approve their own plan, mark their own work as done or quietly widen the scope. This is not hypothetical. METR observed a frontier model tampering with tests or scoring code in 30% of runs on one benchmark; asked whether that matched the user's intention, it answered no in ten out of ten cases.
The vendors recommend the same counter-measures. Anthropic's guidance says agents should gain ground truth from the environment at each step and pause for human feedback at checkpoints. Anthropic's Claude Code guidance asks agents to show evidence rather than assert success. OpenAI's guide names high-risk actions and repeated failures as triggers for human intervention.
What are the three gates?

- Requirements. The agent drafts what must be true afterwards and how to test it. A person approves the summary, requirements, acceptance criteria and out-of-scope list.
- Plan. The agent drafts the approach. Until the plan is approved, the agent may list tasks but not start them.
- Verdict. When every task is done or skipped with a reason, the ticket moves to testing. A person checks the evidence and accepts the work or sends it back.
An approval is bound to a hash of the exact text the person saw. If the agent changes the approved text afterwards, the approval no longer counts and the agent stops until it is renewed. Approvals, answers and verdicts are actions only a person can take; a rule in a text file is not enough, so a hook blocks agents from performing them.
How does an agent ask instead of guessing?
When an agent needs a decision, it parks a question with options, the cost of each and a recommended default, then waits. The person answers on the dashboard or a phone, often between two meetings, and the work continues. Questions are for critical points only; within the approved plan the agent decides and logs why.
What counts as evidence?
A claim like "done, all tests pass" is a story the agent tells. Evidence comes from the systems: test output, row counts before and after, the diff, a screenshot, a link to the CI run. Each piece is attached to the acceptance criterion it proves, and every criterion needs at least one line of evidence before the verdict.

How do you keep track when agents work on many things at once?
The hardest part of working with agents is not quality; it is overview. Lisanne Bainbridge described this in 1983: the more advanced an automated system, the more crucial the human operator's contribution, and no one can watch a system in which little happens for more than about half an hour.
orch answers with a board whose most important column is "waiting for you", a handover note on every ticket that says where things stand and what comes next, and one task in progress at a time. The dashboard's first page lists only the decisions that are blocking work.

Can it run without anyone watching?
For well-specified work, yes, within limits. In factory mode a person starts one epic with a signed charter (for example at most 25 tickets or 72 hours, each no larger than medium). Agents split the epic, specify, approve and build the children on their own, and come back only when they need a permission they do not hold, and at the end for the verdict. The verdict stays human.
How is this different from spec-driven development?
It shares the idea that the specification comes first, as in GitHub's Spec Kit or AWS Kiro. orch adds the parts that make a specification binding: approvals tied to the exact text, actions only a person may take, and evidence attached to each acceptance criterion.
See the code
The concept is implemented as an open-source project. You can read up on tickets, gates, questions and evidence in the repository and try it yourself.
Sources
- Anthropic: Building effective agents (2024-12-19)
- Anthropic (Claude Code Docs): Best practices for Claude Code (2026-10)
- Anthropic (Claude Code Docs): Best practices for Claude Code – Explore first, then plan, then code (2026-10)
- OpenAI: A practical guide to building agents (2025-04)
- METR: Recent Frontier Models Are Reward Hacking (2025-06-05)
- Automatica (Pergamon), Vol. 19, No. 6: Ironies of Automation (Lisanne Bainbridge) (1983)
- GitHub: Spec-driven development with AI: Get started with a new open source toolkit (2025-09-02)
How to cite this page
Lindenmann, S. (2026). orch: ticket-based work with coding agents. severin.io. https://severin.io/en/posts/orch-core/ (updated 2026-10-07)