Method v0.1 • fictional fixtures • local-only collector

Not another agent-maturity quiz. Measure the configuration you actually ship.

A reproducible sample benchmark for coding-agent instruction load, declared MCP surface, permission-risk markers, and explicit review controls. The collector checks known agent configuration files only: no source-code scan, no network call, and no file contents in its output.

Three fictional configurations, one fixed collector

These fixtures contain no real repository, customer, credential, endpoint, or proprietary information. The labels describe comparison profiles, not grades or safety certifications.

Lean / explicit

Minimal repo instruction

94

estimated startup-context tokens

Known config files1
Declared MCP servers0
Risk markers0
Control markers1
Moderate / bounded

Small-team operating note

286

estimated startup-context tokens

Known config files3
Declared MCP servers1
Risk markers0
Control markers5
Manual review

Verbose / open profile

465

estimated startup-context tokens

Known config files3
Declared MCP servers3
Risk markers2
Control markers0

What the collector does — and does not do

Collected

Known agent instruction/config file paths and byte counts; rough instruction-context estimate; declared MCP server count; counts for narrowly defined risk/control text markers; method version and limitations.

Excluded

Application source code, raw prompt or config contents in output, secrets, customer data, runtime tool calls, network traffic, provider billing, hidden system prompts, and claims that a written control is enforced.

Interpretation rule: a marker is a review pointer, not proof. “No marker found” does not mean safe; “control marker found” does not prove enforcement. Exact provider-billed token or cost calculations are outside this method.
python3 coding_agent_hygiene_collector.py /path/to/repo \ --label redacted-team \ --output snapshot.json # Review and redact snapshot.json before external sharing.

Useful review questions after the scan

Context

Which instructions must load every task? Which belong in task-specific docs? Is repeated guidance creating noise without evidence of better outcomes?

Tools

Which MCP servers and shell/browser capabilities are declared? Are production, customer-data, publish, purchase, and destructive actions clearly bounded?

Proof

Do agent-assisted PRs record tests, browser/API evidence, known gaps, human acceptance points, and a rollback note for launch-critical changes?

Turn the snapshot into a founder-readable fix list

A$149–299 Coding Agent Hygiene Review — one public or redacted workflow, returned in 48–72 hours with the local metadata snapshot, configuration map, prioritized context/tool-boundary findings, PR evidence checklist, and a repo-ready operating-note outline. Production access or implementation requires separate explicit authorization.

Free mini-snapshot: one public repo or user-supplied redacted configuration, with 3–5 outside observations. No security certification, compliance opinion, code-quality guarantee, exact token/cost promise, or claim that static configuration proves runtime behavior.