Independent AI Safety Research · Production Evaluation

Precision systems for adaptive intelligence.

Research-grade AI safety tools for measuring behavioral instability, model drift, and safety-relevant regime change during inference.

LLM Evaluation · Inference-Time Stability · RAG & Agentic Systems · MCP Tooling

cos = 0.1524Neutral‑anchored median
0%Antipodal fraction
38.8%L/R coactivation
−1 tokenPre‑leak lead
6 / 6Constraint retained
4 / 6 LR‑action block
Inside the CRAD case study →
Core Thesis
Some failures may have measurable precursors.
Model failure is not always a final answer. Sometimes it is a trajectory.
Click to reveal
TwoQuarks
Testing for measurable precursors to visible failure.
TwoQuarks studies drift, divergence, and candidate pre-critical change during inference. The work translates those measurements into tools for LLM evaluation, monitoring, and operational AI safety.
Click to return

About & Development

Independent AI-safety researcher and evaluation engineer. I build instruments for measuring instability and behavioral drift during inference, and for testing whether those signals precede visible failure.

What I actually do

I'm Jaime Ledesma — independent researcher running TwoQuarks out of Guadalajara. No lab behind me, no cluster. Just the work.

The core research question is whether model failure unfolds as a trajectory rather than a single terminal event, and whether measurable behavioral or internal changes can precede the visible failure. TwoQuarks tests that question without assuming the precursor is universal or causal.

So I build instruments around both views. From the outside — PfV / Molecule, estimating a black-box proxy for structural divergence from API outputs. From the inside — CRAD, measuring state-action divergence in a frozen model while a measured constraint direction remains active. Dopamine is the architecture branch, testing whether representation, integration, memory, and transmission can be exposed as separate mechanisms. Public preprints, code, a Python package, an MCP tool, and a playground make the work inspectable.

who: Jaime Ledesma — independent AI-safety researcher
based: Guadalajara, MX · no lab, no cluster
hypothesis: measurable precursors can precede visible failure
methods: black-box PfV / Molecule · white-box CRAD · experimental architecture Dopamine
ships: public preprints · Dopamine prototype · twoquarks (PyPI) · MCP tool · live playground
stack: Python · PyTorch · HF · APIs · MCP · RAG
status: open to AI-safety collaboration & roles