Quality cannot be the last column on your board
As AI accelerates delivery and evaluators move inside the work, quality needs to connect requirements, tests, defects, evidence, and release decisions from the start.
The work reaches the last column on the board. Then quality assurance begins.
The tester opens the issue and finds a short description. The acceptance conditions were refined in a meeting. The latest design is in a message thread. A dependency changed during implementation. The person who made the decision remembers why, but is now working on something else.
Nothing is necessarily wrong with the output. The problem is that quality has arrived after the context has already left.
That operating model becomes fragile when AI can produce drafts, code, analysis, and campaign assets faster than teams can define, test, and accept them. More work reaches the quality gate, but the gate still depends on people reconstructing what the work was meant to do.
The answer is not a larger final checklist. Quality cannot be the last column on your board. It has to travel with the work.
The evaluator is moving inside the work
On September 18, 2026, Anthropic announced a partnership with Accenture on independent evaluation of frontier AI. The unusual part is not simply that an external organisation will test models. Anthropic says embedded evaluators will have access comparable to employees, allowing them to observe models as they take shape and follow the decisions that govern development and deployment. Read Anthropic's announcement.
The arrangement is specific to frontier AI, and the standards are still being developed. But the operating principle applies much more widely: an evaluator sees more when evaluation happens alongside creation, not only after the finished output is presented.
OpenAI's September 10 Agents API announcement offers another signal. One customer example describes an agent workflow spanning implementation, independent review, remediation, and real-browser validation inside an active repository. Read OpenAI's announcement.
These announcements do not mean every project needs an embedded external evaluator. They show where complex work is heading. Production and evaluation are becoming connected activities rather than separate phases.
For an ordinary product, marketing, operations, or client-delivery team, that changes a practical question. Instead of asking, “Who will check this when it is finished?”, ask, “How will the work produce evidence while it is being done?”
Done is being asked too early
Many boards end with a status called Done. That word often carries more certainty than the project has earned.
It may mean the assignee completed the implementation. It may mean a document was drafted, a page was designed, or an automation was configured. It does not automatically mean the requirement was met, the critical paths were tested, the defects were resolved, the evidence was reviewed, or an accountable owner accepted the release.
AI makes the gap easier to miss because the output can look complete very quickly. A polished result creates pressure to treat verification as a formality. Yet the faster the first version arrives, the more important it becomes to preserve the decisions that explain what should be tested.
A useful definition of done therefore has three connected parts:
- The output exists. The planned work has been produced.
- The evidence supports it. The relevant conditions have been tested and the result is visible.
- The decision is recorded. A named owner has accepted the remaining risk and release state.
If those facts live in different systems, “Done” is still an opinion that has to be reconstructed.
Quality needs work, not just a status
Adding a Review or QA column is useful, but a column only shows location. It does not describe the work required to earn the next transition.
Quality needs its own planned items: test cases, execution runs, defect investigations, regression scope, acceptance checks, and sign-off decisions. Those items need owners, dates, dependencies, and visible outcomes just like implementation work.
In Orbyna Project Management, teams can manage test cases and suites alongside delivery, link tests to requirements or stories, record execution results, and connect failed results to defects. The QA workbench supports defect triage, regression progress, and release-readiness checks.
That changes the shape of the plan. Testing is no longer an unnamed effort squeezed between “finished” and “launch.” It becomes work the team can estimate, assign, sequence, and inspect.
- Create the acceptance checks while the requirement is still being shaped
- Name the owner of each important verification step
- Link test cases to the story or requirement they are proving
- Turn failures into connected defects instead of separate messages
- Keep release sign-off distinct from implementation completion
This does not require every small task to become a heavy test programme. The depth should match the risk. The important change is that quality work becomes explicit before the final handoff.
Connect the chain from requirement to defect
When a test fails, the first question is usually technical: what broke? The more valuable project question is relational: which requirement, dependency, customer promise, or release decision does this failure affect?
That answer is difficult when requirements live in a document, tasks live on a board, tests live in another tool, and defects arrive in chat. The team can count failures but cannot quickly see their consequence.
Orbyna's project record keeps the links visible. Requirements and project knowledge can sit beside issues. Test cases can connect to stories. Defects can link back to failed work. Dependencies show what cannot safely move until the problem is resolved. Issue history preserves how status, ownership, dates, and decisions changed.
The result is an evidence chain:
Requirement → work item → test case → execution result → defect → fix → regression result → release decision
That chain is more useful than a folder of screenshots collected at the end. It allows a product lead to start with a failed test and see what release outcome it threatens. It allows a tester to understand why a condition matters. It allows a future team member to learn why the final decision was reasonable at the time.
Make release readiness a visible decision
A release meeting often combines several different questions:
- Is the planned scope complete?
- Which tests passed, failed, or were not run?
- Are open defects understood and prioritised?
- Did a dependency change the risk?
- Who has authority to accept what remains?
If the answers are assembled during the meeting, the meeting is doing the work of the system.
Use the workflow to make readiness accumulate before the conversation. A practical sequence might be In progress → Ready for test → Testing → Remediation → Ready for sign-off → Released. Transition conditions can require the information your team needs before work moves forward. The exact statuses should fit the project; the key is to keep production, verification, remediation, and acceptance distinct.
Orbyna's project dashboard, board, testing area, QA workbench, and reports show different views of the same delivery record. WIP limits can reveal when verification work is accumulating. Timelines and dependencies show whether a failed condition threatens the date. Workload views help expose when quality depends on one overloaded specialist.
Release readiness then becomes a visible project state, not a confident sentence in a meeting.
Faster production needs faster learning
Embedding quality in the work is not only about preventing bad releases. It shortens the distance between a failure and an improvement to the way the team operates.
Suppose three recent stories reached testing without usable acceptance conditions. Treating each one as an isolated documentation mistake misses the pattern. The workflow itself needs an earlier check.
Or suppose defects repeatedly appear where one service depends on another. The response may not be “test harder.” The project may need better dependency mapping, a reusable regression suite, or a different release sequence.
Orbyna keeps test outcomes, defects, issue history, sprint performance, cycle time, and workflow signals close enough to support that learning. Teams can review where failures entered the process, where remediation waited, and which controls reduced repeat work.
Automation can help with the routine edges: notify an owner when a critical defect is created, add a regression label when affected work changes, or create a follow-up task after a failed condition. But automation should strengthen the evidence chain, not replace the judgement that accepts a release.
The goal is a tighter loop:
Plan → build → verify → remediate → accept → learn
When AI accelerates build, the rest of the loop must become easier to see and improve. Otherwise the organisation produces more first versions while learning at the old speed.
Move quality left without losing the record
“Move quality left” is often understood as testing earlier. That is necessary, but incomplete. An early test that loses its connection to the final decision is still weak evidence.
Start with one active project. Choose a release or customer outcome that matters. Define its acceptance conditions before implementation is complete. Create the important test cases beside the work. Link failures to defects and dependencies. Give sign-off to a named owner. After release, inspect what the evidence revealed about the workflow itself.
The change should make quality more useful, not more ceremonial. Teams should spend less time searching for context, repeating checks, and explaining how a decision was made.
Independent evaluation may be moving inside frontier AI labs, but the broader lesson belongs to every team shipping consequential work: the best time to build the proof is while the work is still taking shape.
Explore Orbyna Project Management to connect delivery, testing, QA, dependencies, and release evidence in one project record, or book a demo around a workflow your team currently verifies at the end.