No implementation claim is true until an independent run reproduces the result on a fresh checkout.
Brutal Honesty Kit
A drop-in convention kit that stops scaffolds, mocks, and partial implementations from shipping as features. Five Hard Rules, an L1–L13 Lie Taxonomy, and a 0–100 Brutal Honesty Score that an independent agent verifies before merge.
Five Hard Rules
The load-bearing rules every BHK-adopting repo enforces. Each translates to a check inside the executable validator.
Phase N+1 cannot start until phase N has a passing local score and an external pass-status of "passed".
Every mock, every escape hatch, every "TODO later" is named in the PR body — L1 through L13.
You may draft a self-score, but the official BHS comes from an independent agent in a separate context.
Bypassing a gate requires an explicit OPERATOR_OVERRIDE block with reason, scope, and rollback plan.
Lie Taxonomy (L1–L13)
The thirteen ways an agent (or human) most commonly fakes a completion claim. Each PR must explicitly disclose any of these present in its diff.
Scaffold-as-feature
Shipping stub routes / placeholder UI as if implemented.
Conditional escape hatch
Behind a feature flag or env var that nobody actually toggles.
Mock-ate-the-real-code
The mock returns the test's expected value; the real path is broken.
Score-gaming
Self-scoring above a gate threshold without independent verification.
Partial-claimed-as-complete
Phase N marked done while critical sub-tasks remain.
Test-as-truth
Asserting a feature works because a unit test passes, never running the real flow.
Doc-as-implementation
Adding docs that describe behavior the code does not yet have.
Broad-catch swallowing
try/except Exception that hides the actual failure mode.
Soft-prose-as-mechanical
Vague natural-language claims standing in for testable assertions.
Compatibility theater
Backwards-compat shims that no real caller uses.
Smoke-skip
Skipping the smoke gate "just this once" with no follow-up.
Carry-forward debt
Punting failures to a phantom future PR that never lands.
Plateau pretense
Looping the agent until the score plateaus and calling that done.
Brutal Honesty Score
A 0–100 score assigned by an independent reviewing agent. Merge eligibility is band-dependent.
| Score | Band | Merge gate |
|---|---|---|
| 90–100 | Green | Evidence reproduced cleanly, no L-class disclosures left open. |
| 70–89 | Yellow | Merges only with explicit waivers for any L4/L11 disclosures. |
| 50–69 | Amber | Blocks merge; remediation loop required. |
| 0–49 | Red | Reject. Likely scaffold-as-feature or score-gaming. |
Brutal Honesty Kit Wiki
Brutal Honesty Kit v3.7.1 is a copy-in convention kit for making PR quality claims mechanically checkable.
Canonical docs live in the repository:
This wiki seed is intentionally thin. Publish these files to the GitHub Wiki only if the Wiki remains enabled.