We use AI to write software faster, behind strict automated checks that decide what may land. Honest Code principles, an honesty gate, and Slop Audit measurement keep the decision space finite so a decidable suite can exhaust it before production.
Firms that turn coding assistants on without changing how code is shaped often get the same result: velocity up, risk up, and a false sense of safety from green test suites. General-purpose models default to class-heavy, state-rich patterns that make exhaustive verification impossible, even when line coverage looks fine.
Generated code can compile, pass shallow tests, and still encode the wrong behavior, a production risk you cannot see in a diff review alone.
Mutable objects and open dispatch create call sequences no suite can finish. Coverage can be green while the real decision space never closes.
I/O mixed into business logic forces tests that agree with themselves. Failures hide until integration or incident.
Asking the model to “be careful” does not remove defect categories. What the shape allows, the model keeps emitting.
Human inspection after generation does not scale. Volume rises faster than the ability to reason about every change.
Line coverage and CI green lights measure execution of lines, not whether the suite can exhaust the decisions that matter.
AI accelerates both good structure and bad. Without a gate, the bad compounds quietly across every merge.
Buckler does not treat AI coding as a hosted chatbot that lands whatever looks plausible. We use assistants inside a build-in-quality discipline: named principles that remove defect categories, an automated honesty gate that refuses dishonest code before it is committed, and Slop Audit, which scores whether the suite can finish the space, plus eighteen enterprise production-readiness dimensions.
A language model on its own is a semi-random generator. At Buckler, we decide what code architecture is allowed and how it must be proved. The Open Honest stack (Honest Code principles, the Honest Framework gate, and Slop Audit) is that control surface.
The point is poka-yoke: make whole classes of bug unbuildable, and keep automated testing able to keep up with the volume of generation.
Engineers use AI inside the repo against stated contracts and data shapes, not as an unchecked author of architecture.
P01–P25 name defect categories to remove up front: open dispatch, Big State, mixed I/O, mocks in the core, and more. Principles tell the model how to avoid them in the first place.
honest-check refuses structurally dishonest code while the AI is coding; honest-test blocks commits that fail the standard. No “mostly honest” merges.
Layer 1 scores whether the codebase can be verified (mutable-state ratio, decision space, determinism). Layer 2 scores the eighteen enterprise dimensions below. Objective measurement, not opinion.
Prompt guardrails constrain what a model is asked to say. The model treats them as suggestions. Buckler adds a different step: when code is produced, mechanical checks examine the code and the tests together, and refuse commits that leave an open decision space or break other named patterns.
Code that is not verify-able is not accepted. Plausible-looking code with no objective checks is what puts the “vibe” in “vibe coding.”
Safe AI-assisted coding is a property of code architecture. Pure cores, I/O only at the boundary, closed dispatch tables, and faults that travel in return values are how we find faults and keep many bug classes from being written at all.
Anyone can generate a passing demo in an afternoon. What matters is whether the same change stays decidable as the system grows, and whether a green CI run means the suite finished the space, or only walked a thin path through an open one.
Beyond Layer 1, Slop Audit Layer 2 scores eighteen production-readiness dimensions. Each has a published threshold, a mechanical inspection procedure, and Present / Partial / Absent scoring. The list below is taken from the open Slop Audit spec quick reference (spec/dimensions/00-quick-reference.md). All eighteen are drafted in v0 form; calibration against the validation set is ongoing work, not a stub list.
Lifecycle labels are the categories the instrument uses (security, data, compliance engineering, ops, architecture, governance, and so on). The one-line threshold is the industry bar the assessor applies.
The Slop Audit instrument claims six framework families. That claim is owned in one place in the open repo (spec/04-compliance-frameworks.md), grounded in the published natural experiment. A dimension citing a framework means its evidence would help an assessor asking about that control. It does not mean the instrument audits for that framework, and a passing score is never a compliance finding by itself.
| Framework | Coverage in the instrument | How to read it |
|---|---|---|
| NIST SP 800-53Documented | 18 of 18 dimensions | Supports evidence for control families such as access control, audit, system and communications protection, and secure development (SA/CM/AU/AC/SC as cited per dimension). |
| SOC 2 (TSC)Documented | 11 of 18 dimensions | Aligns with common Trust Services themes (logical access, change management, monitoring, availability communication) where the dimension files cite them. |
| OWASP (ASVS, API Top 10, SAMM)Documented | 3 of 18 dimensions | Narrow, explicit citations (for example unrestricted resource consumption on rate limiting). |
| DORADocumented | 3 of 18 dimensions | Cited where delivery and ops capabilities map; not a blanket DORA score. |
| CISDocumented | 2 of 18 dimensions | Cited for secrets and container/orchestration benchmarks where the files name them. |
| CNCFDocumented | 2 of 18 dimensions | Cited where cloud-native guidance is relevant in the dimension entries. |
Dimension files also cite other regimes (for example OSFI B-13, ISO/IEC 25010, PCI DSS, GDPR, Quebec Law 25). Those citations exist in the catalog today but are not part of the instrument’s claimed six until a released catalogue maps them dimension by dimension the same way. We do not treat them as certifications Buckler “meets.”
Illustrative alignment (not a documented claim of the instrument): the gated SDLC pattern (spec-before-code, pre-commit honesty checks, CI quality gates) is the kind of control theme you also see in NIST SSDF-style secure development and ISO/SOC-flavored change management. Where a dimension already cites NIST SSDF or SOC 2 CC8.1, that is documented; broader “SDLC maturity” language beyond those citations is illustrative only.
Because generation sits behind principles, an honesty gate, and open measurement of what the suite can finish, the risk posture is simple to state: AI may draft; structure and suite decide what lands; remaining risk is scored in the open.
Constant watch on assistants, models, and editor tooling as they mature.
A tool-agnostic stance: the gate and meter stay; the generator can change.
Structural lint and behavioral checks at commit, not optional style advice.
Pure-function assertions preferred over mocks in the core.
Contracts and data shapes constrain AI input; the gate constrains AI output.
Slop Audit Layer 1 reports whether code can be verified, and fails closed when analysis is unresolved. It does not claim the suite already verified everything.
Layer 2 dimension scores are evidence artifacts for risk and audit conversations, not a substitute for a scoped compliance opinion.
Effective AI-in-coding governance is not a policy binder alone. It is roles, gates, and documents that match how work actually ships. Establishing the discipline early is what keeps AI velocity from becoming unowned debt, and what makes adoption, audit readiness, and delivery the same project.
honest-check / honest-test refuse dishonest structure at edit and commit time. CI carries the same bar. What cannot pass the gate does not merge, regardless of how plausible the model output looked.Principles, honesty gate, and open measurement of what can be proved, not prompt hope.