Spec-Driven Agentic Development
A specification-first workflow can reduce requirements-related rework. Review and tests still need to verify both the specification and its implementation.
You have a feature to build. The requirements are in a Linear ticket. Some context is buried in a Slack thread from two weeks ago. There are notes from a meeting you half-remember. A teammate mentioned an edge case in a standup. Your tech lead has opinions about the API design that live exclusively in their head.
Now you sit down, open Claude Code, and type: "Build the notification system."
The result can miss the requirements when the input is scattered, incomplete or ambiguous. I wrote about this in The Verification Gap: vague input, vague output.
There’s a better way. It’s not a new AI model. It’s a discipline backed by a framework.
Write the spec first. Review it as a team. Freeze it. Then build. And with OpenSpec, this becomes structured, repeatable, and agent-native.
The Scattered Input Problem
Every feature starts as distributed knowledge. Requirements live in different places, in different formats, owned by different people:
- Linear/Jira: The ticket — usually a title and two sentences
- Slack: Three threads, two channels, one DM with the actual decision
- Meetings: Verbal agreements nobody wrote down
- Your head: Technical constraints you know from experience
- Someone else’s head: Edge cases, business rules, compliance requirements
- Figma/Docs: Design mockups that may or may not match the latest decisions
This is normal. Requirements should emerge from conversations and collaboration. The problem isn’t that the information is scattered — it’s that we skip the step of assembling it into a coherent specification before we start building.
Instead, we go straight from scattered inputs to code. We treat the AI agent like a mind reader. And then we spend three days iterating on code that should never have been written in the first place — because the foundation was wrong.
Two Approaches
The traditional approach:
Scattered info → Vague prompt → Code → Review → "That's not what I meant" → Rewrite → RepeatYou iterate on code. Every cycle is expensive: the AI rewrites files, tests break, reviewers re-read everything, context shifts. Three rounds of code review later, someone realizes the API contract was wrong from the start.
The spec-driven approach:
Scattered info → Spec draft → Review spec → Iterate spec → Approve spec → Implement → Test and review → Revise as neededReviewing requirements before implementation can expose ambiguity while changes are still relatively cheap. A reviewed specification improves the input to the agent; it does not guarantee that the first implementation is correct.
The shift is to begin reviewing requirements before implementation: move the iteration from code to spec.
Enter OpenSpec
OpenSpec is an open-source spec framework designed as a planning layer for AI coding agents. It takes the spec-driven philosophy and gives it structure, repeatability, and direct agent integration.
What it provides:
- Structured artifacts — proposal, design, specs, and tasks as separate Markdown files
- Capability-based organization — specs organized by what the system does, not by ticket number
- Change management — each change proposal lives in its own directory with all context
- Agent commands —
/opsx:propose,/opsx:apply,/opsx:archiveintegrated into coding agents - Git-native — everything lives in the repo, version-controlled alongside the code
- Agent-agnostic — works with 30+ coding agents (Claude Code, Cursor, GitHub Copilot, Codex, and more)
No API keys. No external dependencies. Just structured Markdown files in your repo.
The Directory Structure
openspec/
├── specs/ # Living capability specs (persistent)
│ ├── notification-delivery/
│ │ └── spec.md # Maintained intended behavior
│ ├── user-permissions/
│ │ └── spec.md
│ └── payment-processing/
│ └── spec.md
├── changes/ # Active change proposals
│ ├── add-notification-batching/
│ │ ├── .openspec.yaml # Change metadata
│ │ ├── proposal.md # What & why
│ │ ├── design.md # How (technical decisions)
│ │ ├── tasks.md # Implementation checklist
│ │ └── specs/
│ │ └── notification-delivery/
│ │ └── spec.md # Spec delta (what changes)
│ └── archive/ # Completed changes
│ └── 2026-03-26-add-notification-batching/
│ └── ...Three directories, three purposes:
openspec/changes/<name>/— Active proposals. Each change gets its own directory with four artifacts. This is where spec work lives during proposal and review.openspec/specs/— Maintained specifications. Describe intended behavior for the current version. Organized by capability, not by ticket. Populated via/opsx:archivewhen changes are completed.openspec/changes/archive/— Completed changes, moved here with a date prefix. Full history of proposals, designs, and tasks alongside the code that implemented them.
The separation is deliberate. Active work goes in changes/, maintained specifications live in specs/, and the archive preserves history. Readers can use openspec/specs/ to understand intended behavior, then check the implementation and test evidence for agreement.
The Four Artifacts
Every change proposal in OpenSpec produces four files. Each serves a distinct purpose, and together they form a complete, reviewable description of a change before any code is written.
1. Proposal (proposal.md) — The Why
The notification service below is an illustrative design example. Its windows and requirements are example choices to review, not telemetry from a private system.
## Why
The notification system sends emails one-at-a-time. At scale, this creates
delivery delays of 30+ minutes during peak hours. Users report receiving
notifications long after the triggering event, reducing trust in the platform.
## What Changes
- Add a batching layer that groups notifications by recipient and channel
- Implement configurable batch windows (default: 5 minutes for email, immediate for push)
- Add a digest template that combines multiple notifications into a single email
- Update the delivery pipeline to route through the batch queue
## Capabilities
### Modified Capabilities
- `notification-delivery`: Add batching and digest support
### New Capabilities
- `notification-batching`: Configurable batch windows and digest generation
## Impact
- Delivery pipeline: new batch queue before email dispatch
- Database: new batch_config table, notification_batch junction table
- Templates: new digest email template
- API: new endpoint for batch configurationThe proposal captures what and why without prescribing how. The “Capabilities” section creates the contract between proposal and specs — it names exactly which capabilities are new or modified.
2. Design (design.md) — The How
Technical decisions with rationale. This is where you make choices explicit and debatable before anyone writes code.
## Decisions
### D1: Durable pending notifications with in-process batching
Persist each accepted notification in a durable pending-notification table
before acknowledging acceptance. An in-process accumulator groups pending
rows; the database remains the source of truth. On startup, reload pending
rows and resume expired batches. Record delivery acknowledgement separately.
Use an idempotency key with the delivery provider where supported and account
for possible duplicate delivery after a crash between send and acknowledgement.
### D2: Windows per channel; accumulation per recipient and channel
Email: 5 minutes. Push: immediate. SMS: 15 minutes.
Configure windows by channel, but keep a separate accumulator for each
(recipient_id, channel) pair. Never combine different recipients in a digest.
### D3: Digest contains a summary and deep links
Summarize the included notifications with links to details. Validate usability
with the product team; this example claims no benchmarked click-through gain.
## Risks / Trade-offs
- Batching deliberately adds delivery latency
- Durable recovery requires retries, acknowledgement and duplicate handling
- Graceful shutdown should flush when possible; crashes recover from storage
- Provider idempotency and retention limits must be checked explicitlyEach decision has a number (D1, D2, D3) so reviews can reference them precisely: "I disagree with D1 because..." instead of "somewhere in the design you said...".
3. Spec (spec.md) — The Contract
Requirements in SHALL format. Scenarios in GIVEN/WHEN/THEN. This is the frozen contract the implementation must satisfy.
## Purpose
Group durable pending notifications by recipient and channel into digests.
## Requirements
Requirement: The system SHALL persist an accepted notification before
acknowledging acceptance and retain it until delivery is acknowledged.
Requirement: Accumulation SHALL remain separate per recipient and channel.
Requirement: Email and SMS SHALL use configured batch windows. Push SHALL
bypass the accumulation delay, while using the same durable delivery tracking.
Requirement: A digest SHALL summarize its included notifications with deep links.
Requirement: Startup SHALL recover pending rows and dispatch expired batches.
Requirement: Graceful shutdown SHALL attempt a flush; unacknowledged rows
SHALL remain durable for recovery. Shutdown alone is not a no-loss guarantee.
Requirement: Delivery SHALL use stable idempotency keys where supported and
explicitly handle retries and the send-before-acknowledgement crash window.
## Scenarios
Scenario: Email batching within window
GIVEN one recipient receives A at T+0 and B at T+2min
WHEN its 5-minute window expires
THEN one digest includes A and B
Scenario: Recipient isolation
GIVEN two recipients receive notifications in the same channel window
THEN their notifications remain in separate batches
Scenario: Crash recovery
GIVEN accepted notifications remain unacknowledged in durable storage
WHEN the process restarts
THEN it reloads them and retries eligible delivery without discarding them
Scenario: Single notification
GIVEN a batch contains only one notification
WHEN the window expires
THEN a standard message is sent instead of a digest
Scenario: Send succeeds before process crash
GIVEN delivery occurred but acknowledgement was not persisted
WHEN recovery retries
THEN the stable idempotency key or documented duplicate-handling policy appliesScenarios provide concrete acceptance criteria that can be translated into tests. Reviewers should check that those tests represent the intended behavior and that important cases are not missing. Passing them supplies evidence within their coverage.
4. Tasks (tasks.md) — The Plan
Checkboxes. Ordered. Each task is a unit of work the agent can execute independently.
## Tasks
### Persistence and batching
- [ ] Persist accepted notifications with recipient, channel and stable delivery key
- [ ] Accumulate separately for each recipient/channel pair
- [ ] Configure channel windows (email: 300s, push: 0, sms: 900s)
- [ ] Recover pending rows on startup and dispatch expired batches
- [ ] Retain unacknowledged rows on crash or unsuccessful shutdown flush
- [ ] Implement acknowledgement, retries and duplicate-delivery handling
### Digests and configuration
- [ ] Build summaries with counts and deep links
- [ ] Send a standard message for a single-notification batch
- [ ] Add batch configuration storage and an authorized configuration endpoint
### Tests
- [ ] Two notifications within one window produce the expected digest
- [ ] Different recipients never share a batch
- [ ] Push bypasses accumulation delay
- [ ] Graceful shutdown and abrupt restart retain pending work
- [ ] Crash after send and before acknowledgement exercises duplicate handling
- [ ] Existing notification behavior remains coveredWhen you run /opsx:apply, the agent uses the design, specification and task list to guide implementation. It may still infer missing details or make incorrect choices; review those decisions and verify the resulting code.
The Workflow
Step 1: Propose
/opsx:propose add-notification-batchingThe agent gathers context from your codebase, existing specs, and any input you provide (ticket, Slack thread, meeting notes), then generates all four artifacts. You can also write them manually — OpenSpec doesn’t force you through a CLI.
This is the synthesis step: the agent turns supplied context into a structured proposal. Review it for missing information, unresolved decisions and assumptions that need to be made explicit.
Step 2: Draft PR (Spec Only)
Create a pull request containing only the spec artifacts. No implementation code. Mark it as draft.
git checkout -b spec/notification-batching
git add openspec/changes/add-notification-batching/
git commit -m "RFC: Notification batching"
gh pr create --title "RFC: Notification batching" --draftThe RFC: prefix makes it instantly clear: this is a spec review, not a code review. Different reviewers, different expectations, different review criteria.
This is the most reviewable PR your team will ever see. A reviewer doesn’t need to understand code paths, trace execution flows, or run tests. They read a document and answer one question: "Is this what we should build?"
Step 3: Multi-Pass Review
This is where both humans and agents provide input. Multiple eyes, multiple passes:
Human reviewers catch:
- Business logic errors (“We also need to handle enterprise tier rate limits differently”)
- Missing requirements (“What about GDPR? Users need to opt out of digest emails”)
- Scope creep (“Let’s not build the analytics dashboard in v1”)
- Organizational context the agent can’t know
AI agent reviewers catch:
- Technical inconsistencies (“Design D2 contradicts requirement 3”)
- Missing edge cases (“What happens if the batch queue is full on restart?”)
- Security gaps (“The batch config endpoint needs authentication”)
- Numbered decision references make feedback precise
You can even have the agent review the spec from a different angle:
Review openspec/changes/add-notification-batching/ as a senior backend engineer.
Focus on: scalability, failure modes, and operational concerns.
What's missing? What will break at scale? What will wake someone up at 3am?Different perspectives can surface different gaps. Review time depends on the decisions involved; use each pass to resolve specific questions rather than assuming a fixed speed advantage.
Step 4: Freeze
When the team approves the PR, merge it. The spec is now frozen. It’s the contract. All scope decisions and edge case handling are locked. If requirements change later, you update the spec first. The spec leads, the code follows.
Step 5: Implement
Create a new branch and PR for implementation:
/opsx:applyThe agent reads the proposal, design, specs and tasks as explicit implementation context. That is a stronger starting point than an underspecified feature request, but ambiguity and mistaken inferences can remain. Test and review both the specification and implementation, then revise as needed.
Step 6: Archive
After the implementation merges:
/opsx:archiveThis does two things:
- Syncs spec deltas from
openspec/changes/<name>/specs/intoopenspec/specs/— updating the maintained capability documents with the approved specification changes - Moves the change to
openspec/changes/archive/YYYY-MM-DD-<name>/— preserving the full proposal, design, and task history
The result is an updated set of maintained specifications and an archive of the changes processed through this workflow. Keeping those documents aligned with the running system still requires review, tests and disciplined maintenance.
Why This Works
Structure Reduces Ambiguity
The four-artifact format gives reviewers a place to examine the why (proposal), how (design), requirements (spec) and plan (tasks). That structure can expose omissions; it does not force semantic completeness. Agents may still infer missing details, so review both the artifacts and implementation.
Capability-Based Organization Scales
Organizing specs by capability (notification-delivery, user-permissions) instead of by ticket (ENG-123, JIRA-123) means specs outlive sprints. When a new developer joins or a new ticket references notification behavior, they find openspec/specs/notification-delivery/spec.md — not a six-month-old ticket buried in Linear.
Scenarios Are Executable Verification
GIVEN/WHEN/THEN scenarios can become test cases. Review whether tests cover the intended behavior and whether the specification itself is correct. Passing scenarios provide evidence within their coverage; they do not close every verification gap.
Iteration is Cheap
Correcting requirements before implementation can avoid downstream rework. Assess the complete effort needed to resolve the issue.
Alignment Happens Before Execution
Misunderstood requirements can be costly. Spec review gives the team an opportunity to resolve disagreements before implementation.
Agents Get Better Input
A clear, reviewed specification can reduce requirements-related rework. Agents can still misunderstand it, implement it incorrectly or miss cases that the document never considered.
Git Makes Specification Changes Reviewable
Keeping specifications beside code makes their history and proposed changes easier to review. Version control records changes, but tests and review must still check semantic agreement between specification and implementation.
The Spec Is Documentation
When the feature ships, the spec stays. Six months from now, when someone asks “Why does the notification system batch emails every 5 minutes?” the answer is in openspec/specs/notification-delivery/spec.md, with the archived proposal that explains the reasoning.
Common Objections
"This slows us down"
Specification writing and review take time. They can reduce rework caused by misunderstood requirements, but the net effect depends on the project; no universal time-saving percentage is established here.
"Requirements change constantly"
Update the specification as requirements change, then review the corresponding implementation changes. This makes the intended change visible and reviewable.
"Isn't this just waterfall?"
Keep each proposal scoped to one change and use short feedback cycles. Document length and review time depend on the work.
"My agent is smart enough to figure it out"
An agent may misunderstand even a detailed description. Specifications make the intended behavior easier to examine, but the team must still check whether it meets the actual need and whether the code implements it.
"The spec will get outdated"
The archive workflow can carry approved specification deltas into the maintained documents. That synchronization is a document operation, not proof that the code implements the resulting specification. Review and tests must establish that separately.
Setting It Up
Installation
npm install -g @fission-ai/openspec@latest
cd your-project
openspec initThis creates the openspec/ directory structure and configures your agent. Works with Claude Code, Cursor, GitHub Copilot, and 30+ other tools out of the box.
Agent Instructions (CLAUDE.md)
Add the spec-first rule to your project’s CLAUDE.md:
### Spec-First for Complex Changes
Any ticket that touches core business logic MUST follow spec-first development.
Process:
1. /opsx:propose — Generate proposal, design, spec, and tasks
2. Draft PR — Spec-only, assign reviewers
3. Iterate on spec feedback
4. Freeze — Merge spec PR
5. /opsx:apply — Implement against frozen spec
6. Test and review the implementation; revise the spec or code as needed
7. /opsx:archive — Sync approved spec changes to maintained docs, archive changePR Naming Convention
# Spec PRs (draft until frozen)
RFC: Notification batching
RFC: User permission redesign
# Implementation PRs (reference the spec)
feat: Add notification batching (openspec/changes/add-notification-batching)A Spec-Writer Skill
Create .claude/skills/spec/SKILL.md with the following content. The first line of the file must be the frontmatter delimiter.
---
name: spec
description: Create a spec-first change proposal using OpenSpec
disable-model-invocation: true
argument-hint: "[change-name]"
---
Run /opsx:propose for: $ARGUMENTS
After generating the artifacts, review them yourself and ask me
for any missing context before I create the draft PR.
Ensure the proposal has:
- Clear problem statement (why)
- Explicit capability mapping
- Numbered design decisions
- SHALL requirements with GIVEN/WHEN/THEN scenarios
- Ordered task checklistNow /spec add-notification-batching triggers the entire spec creation workflow.
Closing the Verification Gap
In The Verification Gap, I argued that the biggest risk in AI-assisted development isn’t bad code — it’s the inability to verify that what the agent built is what you actually wanted.
OpenSpec attacks that gap with surgical precision. The spec is the intention, made explicit. SHALL requirements become assertions. GIVEN/WHEN/THEN scenarios become test cases. Design decisions are numbered and reviewable. The gap between intention and implementation shrinks from “Did the agent understand what I was thinking?” to “Does the code match the spec?” — and that’s a much easier question to answer.
Spec-driven development combines structured requirements with review. Its value is making assumptions and acceptance criteria explicit. Teams should assess whether it reduces their own rework instead of treating faster delivery as a guaranteed outcome.
Better input, better output. It was always that simple.