⧁
Alex Price / StellarRequiem
Security Research · AI-Systems Verification · Responsible Practice
Professional Summary
I build the trust layer for AI-era work. As models and agents produce more claims than anyone can check, the dangerous failures are the confident-sounding results nobody verified. One rule governs everything I deliver: no belief without verification — every result ships with evidence a third party can independently re-run. I apply the same discipline to security: offensive technique under explicit authorization, with a complete, tamper-evident record of what was tested and found.
Core Competencies
Vulnerability research & PoC development
Verification & statistical-validity engineering
Authorization-gated security testing
AI-infrastructure / MCP security
Coordinated / responsible disclosure
Append-only, hash-chained audit systems
Red-team & offensive tooling development
Agentic-system evaluation & calibration
Selected Work — public, CI-verified · github.com/StellarRequiem
scope-gate Responsible security
A deny-by-default authorization gate — test only what you are explicitly authorized to test. Ships with a responsible-research charter. The boundary that makes dual-use work safe.
mcp-assure MCP tool gate
Deny-by-default policy + hash-chained receipts for agent MCP tool calls. A runtime gate you can re-run — not a full SOC, not “stops all attacks.” Proof surface: xclusivexo.com/mcp-assurance/.
agent-control stack Local control planes
Session-scoped browser/desktop/CUA through AdaptiveGate — agent-control, browser-leash, desktop-leash, agent-soc. Local-first and arm-gated; not ambient OS takeover and not an enterprise SIEM.
mcp-bench Security benchmark
An independent, reproducible benchmark of whether MCP security scanners catch authorization-logic bugs — seeded with real, responsibly-confirmed findings. Two mature SAST scanners catch the control bugs but miss the authz-logic class.
verity-core The gate
Refuses an accuracy claim until it clears statistical hygiene — sample size, out-of-sample, leakage, and lift over the base rate — then proves it: a claim ships a re-runnable command and the number must reproduce or CI fails. 17 domain packs, a CI gate, and an MCP tool.
verified-ai-labor The platform
A validated organization of agents routing work through a 13-stage pipeline — every result-claim verity-gated, every action hash-chain-logged, observable in a live console.
calibration-log The proof
A public, hash-chained prediction record scored over time (Brier + calibration) — honesty that cannot be doctored after the fact. Reports the real number whether or not there is an edge.
scorecheck The adjudicator
Adjudicates a published benchmark claim against its raw run-logs — REPRODUCED / DID-NOT-REPRODUCE / CHERRY-PICKED — sealed into a re-runnable receipt. Surfaces the dropped, flipped, and fabricated rows that re-run leaderboards and reproducibility badges miss.
groundtruth-bench Citation-faithfulness benchmark
A cryptographically committed corpus scored offline to a byte-identical hash across machines — where RAG eval (RAGAS/ARES) is online and uncommittable. Reports where the scorer fails, not just the flattering number.
Approach & Standard
- Verified work, or it doesn't ship. Public repositories carry their own tests; the CI badge is the claim.
- Runnable proofs. Every result is accompanied by the exact command needed to reproduce it.
- Authorized & audited. Security work is performed only within explicit scope, with an append-only record of activity.
- Honest gaps, stated. Every deliverable names what it did not verify. An unverifiable claim does not count.
Credentials & References
Verifiable identity attestation, professional certifications, security-platform researcher profiles, and references available on request.
Verified work, or it doesn't ship. · no belief without verification
Alex Price · StellarRequiem