Precision systems for adaptive intelligence.
Research-grade AI safety tools for measuring behavioral instability, model drift, and safety-relevant regime change during inference.
LLM Evaluation · Inference-Time Stability · RAG & Agentic Systems · MCP Tooling
Research depth. Production readiness.
TwoQuarks is organized as a portfolio of AI safety research, instruments, framework design, and development work. Each layer points to evidence, code, or operational tooling.
Research
Preprints, empirical results, cross-model findings, statistical controls, and reproducible safety probes.
→ White-boxCRAD
Constraint-Retained Action Divergence: neutral-anchored detection of pre-leak state–action dissociation in a frozen model. Most active research front.
→ Active prototype Model architectureDopamine
An experimental small language model exploring multi-field encoding, shared latent integration, memory pathways, and selective output transmission.
→ ToolingInstruments
Molecule, PyPI package, MCP integration, provider adapters, and operational instability analysis.
→ ArchitectureFramework
The internal TwoQuarks structure: PfV, ΔL3, six flavors, inference-time control, and behavioral regime mapping.
→ Live InteractivePlayground
Run the safety probes yourself: live Molecule-style response divergence across the C1–C5 behavioral regimes.
→ ProfileAbout & Development
Independent AI safety research, engineering background, open-source work, resume, GitHub, and contact.
↓About & Development
Independent AI-safety researcher and evaluation engineer. I build instruments for measuring instability and behavioral drift during inference, and for testing whether those signals precede visible failure.
What I actually do
I'm Jaime Ledesma — independent researcher running TwoQuarks out of Guadalajara. No lab behind me, no cluster. Just the work.
The core research question is whether model failure unfolds as a trajectory rather than a single terminal event, and whether measurable behavioral or internal changes can precede the visible failure. TwoQuarks tests that question without assuming the precursor is universal or causal.
So I build instruments around both views. From the outside — PfV / Molecule, estimating a black-box proxy for structural divergence from API outputs. From the inside — CRAD, measuring state-action divergence in a frozen model while a measured constraint direction remains active. Dopamine is the architecture branch, testing whether representation, integration, memory, and transmission can be exposed as separate mechanisms. Public preprints, code, a Python package, an MCP tool, and a playground make the work inspectable.