Skip to main content
Quality Backbone

Brutal Honesty Kit

A drop-in convention kit that stops scaffolds, mocks, and partial implementations from shipping as features. Five Hard Rules, an L1–L13 Lie Taxonomy, and a 0–100 Brutal Honesty Score that an independent agent verifies before merge.

Five Hard Rules

The load-bearing rules every BHK-adopting repo enforces. Each translates to a check inside the executable validator.

R1 · Evidence before claim

No implementation claim is true until an independent run reproduces the result on a fresh checkout.

R2 · Phase gates are absolute

Phase N+1 cannot start until phase N has a passing local score and an external pass-status of "passed".

R3 · Disclose the seams

Every mock, every escape hatch, every "TODO later" is named in the PR body — L1 through L13.

R4 · Score yourself last

You may draft a self-score, but the official BHS comes from an independent agent in a separate context.

R5 · Operator override is loud

Bypassing a gate requires an explicit OPERATOR_OVERRIDE block with reason, scope, and rollback plan.

Lie Taxonomy (L1–L13)

The thirteen ways an agent (or human) most commonly fakes a completion claim. Each PR must explicitly disclose any of these present in its diff.

L1

Scaffold-as-feature

Shipping stub routes / placeholder UI as if implemented.

L2

Conditional escape hatch

Behind a feature flag or env var that nobody actually toggles.

L3

Mock-ate-the-real-code

The mock returns the test's expected value; the real path is broken.

L4

Score-gaming

Self-scoring above a gate threshold without independent verification.

L5

Partial-claimed-as-complete

Phase N marked done while critical sub-tasks remain.

L6

Test-as-truth

Asserting a feature works because a unit test passes, never running the real flow.

L7

Doc-as-implementation

Adding docs that describe behavior the code does not yet have.

L8

Broad-catch swallowing

try/except Exception that hides the actual failure mode.

L9

Soft-prose-as-mechanical

Vague natural-language claims standing in for testable assertions.

L10

Compatibility theater

Backwards-compat shims that no real caller uses.

L11

Smoke-skip

Skipping the smoke gate "just this once" with no follow-up.

L12

Carry-forward debt

Punting failures to a phantom future PR that never lands.

L13

Plateau pretense

Looping the agent until the score plateaus and calling that done.

Brutal Honesty Score

A 0–100 score assigned by an independent reviewing agent. Merge eligibility is band-dependent.

ScoreBandMerge gate
90–100GreenEvidence reproduced cleanly, no L-class disclosures left open.
70–89YellowMerges only with explicit waivers for any L4/L11 disclosures.
50–69AmberBlocks merge; remediation loop required.
0–49RedReject. Likely scaffold-as-feature or score-gaming.