Case Study — HLE-AIO
A full pre-release security and quality audit of a real application, start to finish: static analysis, a four-agent security review, a multi-agent QA sweep of every screen, and how each finding was handled.
The best argument for a development method is a worked example you can inspect. HLE-AIO (HLE All-in-One, source) is a self-hosted family-management application — one container, one login, a dozen modules covering finance, health, meals, home care, travel, files, documents, and more. It's built entirely under VibeSecOps, and before its open-source release it went through a full audit. This page is that audit.
The point of this case study is not "the audit found nothing." It found several real issues — that's what a real audit does. The point is the method: layered automated and AI-assisted review, honest classification of what turned up, and every security-relevant finding remediated (with a regression test) before release. Honesty about the gaps is what makes the parts that passed credible.
The audit in layers
The audit ran in four layers, cheapest and broadest first, most expensive and deepest last.
Layer 1 — Static analysis and the standard gates
The CI guardrails run on every change, but the audit re-ran them from a clean slate:
- Type-check, lint, and the unit-test suite: all green.
- Static security scan (security-audit, secrets, SQL-injection, and framework rulesets): zero findings on the application source — a direct result of rule 3 (parameterized SQL only) and rule 7 (no unsafe sinks) being enforced from the start rather than retrofitted.
- Dependency audit: a handful of transitive-dependency advisories, triaged and resolved by version bumps — the ordinary maintenance surface of any Node/Bun project, none reachable in a way that affected users.
Layer 2 — Focused security review, four parallel reviewers
Four independent reviews ran in parallel, each with a specific charter: the authentication and session surface; file storage, media, and public sharing; per-module authorization and injection; and deployment, container, and open-source hygiene. Splitting the surface this way means each reviewer goes deep instead of skimming everything.
What the security review confirmed was already right is the larger part of the story:
- Every one of the application's server functions carries its auth middleware; tenant scoping is applied consistently on reads and on update/delete.
- All SQL is parameterized; array handling and full-text search use bound parameters; a spreadsheet-formula-injection guard protects CSV exports.
- Rich-text and markdown rendering paths are safe — unsafe HTML sinks are banned in CI and the render paths build DOM without them.
- First-run setup is race-free (an atomic advisory-lock + conditional insert), so a freshly deployed public instance can't be hijacked by a stranger racing to claim the admin account.
- Session revocation, last-owner guards, and secret-stripping on API responses all held up.
And what it caught, with each finding's disposition:
| Finding | Severity | Disposition |
|---|---|---|
| A document-serving endpoint returned uploaded files inline without forcing a safe content type — a stored-XSS vector. Found independently by two of the four reviewers. | High | Remediated before release — ported the exact guard the file module already used. |
| A public share link's download cap could be bypassed with HTTP range requests. | Medium | Remediated. |
| A server-side request to a user-supplied integration URL wasn't restricted from internal addresses (SSRF surface). | Low/Med | Remediated with an allowlist + private-range block. |
| Login response timing could distinguish real accounts from unknown ones (enumeration). | Medium | Remediated — constant-time path for every outcome. |
| Session tokens stored in the database in plaintext rather than hashed. | Low | Remediated — hash at rest, raw token only in the cookie. |
That the same stored-XSS issue was surfaced by two independent reviewers is exactly why the review runs in parallel with overlapping charters. Redundancy across reviewers catches what any single pass can rationalize away.
Layer 3 — Full-surface QA sweep, every screen
An application isn't done when it compiles — every button has to actually do something. A fleet of QA agents drove a real browser through every screen in the app, module by module: load each page, watch for console errors and failed requests, then exercise real create/read/update/delete flows, file uploads, and multi-step wizards, verifying each mutation persisted.
Across roughly a hundred and thirty screens, the sweep confirmed the important things: no broken authorization redirects, no injection, no dead core flows. Balance math, cascade deletes, validation messages, and confirm dialogs behaved. It also caught real bugs — which is the entire reason you run it before release rather than letting users find them:
- A rich-text editor that crashed on input. The wiki page editor threw an unhandled error on the first keystroke — a component re-render loop — dropping the user into an error boundary. Editing that module was effectively broken. Caught by the sweep, not by a customer. Fixed before release.
- A data-encoding bug in the same module that double-serialized stored content, causing pages to render raw markup. Fixed with a regression test.
- A destructive action missing its confirmation dialog (deleting a single transaction), inconsistent with every other destructive action in the app — caught by comparing a module against its own norm.
- An admin backup endpoint returning an unfriendly raw error instead of a handled message when an optional system tool was absent from the host.
- A feature with server support but no UI wiring (file tagging), and a recurring accessibility-semantic console warning surfaced consistently by three independent agents and fixed in one shared component.
Each was logged with severity, exact location, and reproduction, then triaged fix-now versus track-for-later. The security-relevant and release-blocking ones were fixed before ship; the cosmetic ones were tracked.
This is the honest part. A full-surface sweep of a real application found a module-breaking crash and a data bug — and that's a success, because it found them pre-release, with reproductions, in one automated pass. A method that claims it never finds anything isn't being run seriously.
Layer 4 — Human verification and sign-off
Every finding above was read, reproduced, and dispositioned by a human — not taken on the reviewers' word. Security-relevant fixes each got a regression test that fails on the pre-fix code, per the rulebook. "The AI said it's fixed" is not sign-off; a failing-then-passing test and a human who understands the change is.
What this demonstrates about VibeSecOps
- The invariants paid off. The zero-finding static security scan and the clean authorization sweep aren't luck — they're what happens when tenant scoping, parameterized SQL, and banned unsafe sinks are enforced from the first commit instead of audited in at the end.
- The method still finds things. No process makes a real application perfect. What VibeSecOps guarantees is that the things it finds are found before release, classified honestly, and fixed with a test that keeps them fixed.
- AI scaled the review, humans owned it. Multiple reviewers and a full-surface QA sweep at this depth would be impractical to do by hand on every release. AI made the breadth affordable; the human verification and sign-off made it trustworthy. That division — AI drafts and scales, the human verifies and owns — is the whole method in one audit.