// premise
Is the drug target real?
I build open-source computational biology pipelines that interrogate whether a drug target is actually real — and try as hard to falsify a hypothesis as to confirm it. The methods span human genetics and GWAS interpretation, single-cell transcriptomics, structure-based virtual screening and cheminformatics, and protein-language-model benchmarking. Every project ships a pre-registered hypothesis, content-addressed (SHA-locked) verdicts so results can't quietly drift, and null results published next to the positive ones. Most of it runs end-to-end on free Kaggle GPUs.
pre-registered hypotheses · content-addressed verdicts · gaps documented, not smoothed · free compute
// now
What I'm working on
Current, active lines of work.
Duchenne muscular dystrophy — what a reading-frame rule can and cannot predict
A layered ablation of rule-based models for therapeutic exon-skipping targets in DMD, built from Ensembl coordinates and ClinVar. Started as modulo-3 arithmetic; a reader whose child carries a del3–7 deletion found the blind spot that the domain filter created, and the correction is now the scientific content: crisprking/dmd-exon-skip-ablation.
Nav1.7 (SCN9A) — the genetic–pharmacological asymmetry of a “perfect” pain target
Why lifelong genetic loss of Nav1.7 abolishes pain at no cardiovascular cost, yet acute pharmacological block keeps failing in the clinic — biophysics, expression, and human-genetics arms converging on a two-sided constraint: crisprking/nav17-asymmetry.
Medical school — MS1, UAG International MD
Training clinically while keeping an active open-source computational research line.
Open drug-target auditing
Extending falsifiable-targets — pre-registered, SHA-locked verdicts on whether a target's genetics actually support the proposed direction of effect.
// selected work
Reliability & audit frameworks
Reusable methods for deciding whether a target — or a model’s or metric’s claim about it — can actually be trusted.
A content-addressed audit engine for drug-target direction of effect — inhibit versus activate — built on Open Targets colocalization and benchmarked against approved drug–target pairs. Rules are SHA-locked so a verdict cannot silently change, and the engine is allowed to refuse when the genetics are ambiguous. The production layer adds batch auditing, Nextflow and Snakemake DAGs, and a Docker image.
known gaps: 3 documented failure modesTests whether gnomAD LOEUF constraint works as a drug-target safety oracle, as it is routinely used. It does not: LOEUF remains a solid predictor of genetic loss-of-function tolerance, the thing it was built to measure, but is near-useless for predicting clinical safety of inhibition. A negative result, shipped as one.
A one-page, reproducible target-confidence summary that resolves the underlying evidence into an explicit verdict — including refusal — rather than a score whose provenance is unrecoverable.
When can you trust ESM-C zero-shot variant-effect rankings, and when can you not? A reproducible protein-language-model benchmark showing that a model’s self-consistency across sizes does not predict its accuracy, plus a calibration tool for deciding when to rely on it.
One fixed nine-stage QSAR and conformal-prediction pipeline — curation, noise-floor estimation, random-forest QSAR on ECFP4, scaffold-split cross-validation, conformal prediction, applicability domain, freedom-to-operate, design, synthesis triage — run unchanged across ENPP1, NLRP3, TYK2 and IRAK4, spanning a 23× range in public data. Its most valuable outputs are the two refusals.
Target discovery & mechanism
Public data to a defensible shortlist or a mechanistic claim, end-to-end, on free compute.
Why the best genetically validated pain target keeps failing in the clinic. Argues a genetic–pharmacological asymmetry: lifelong genetic loss of Nav1.7 is cardiovascular-silent, while acute pharmacological block causes on-target autonomic toxicity. Single-cell atlas analysis, a leaky homeostatic-compensation model, and gnomAD / ClinVar / Open Targets human genetics.
A structure-based paralog selectivity counter-screen for the immuno-oncology target ENPP1: dock candidates into ENPP1 and its close relatives ENPP2 (autotaxin) and ENPP3 in one identical zinc-centered box, then rank by cross-paralog margin rather than raw affinity. The top affinity binder reversed to an off-target liability. Ships native controls, bootstrap CIs and two honest negative benchmarks; runs on a free Kaggle T4.
Maps type 1 diabetes GWAS loci to the pancreatic cell types they likely act in, using τ-based cell-type specificity across the HPAP single-cell RNA-seq atlas, cross-validated against autoantibody-positive pre-clinical transcriptional change.
Runs a standard ChEMBL-driven target-discovery pipeline on Madurella mycetomatis, a neglected fungal pathogen nobody had mapped — and catches the pipeline’s own artifact: the top-ranked gene was 384 duplicate records of HDAC4. Triaged to a hardened shortlist, with the audit that caught the error shipped alongside.
An eight-stage open drug-discovery pipeline for Chagas disease targeting the T. cruzi cysteine protease cruzain: three parallel scoring tracks fused into a consensus, a selectivity counter-screen against three human cathepsins, and ADMET filtering. Runs on a free Kaggle T4.
An empirical evaluation of which public-data features actually predict PROTAC E3-ligase tractability — and which are just correlated with how well-studied a ligase is.
Genomics & variant interpretation
Rare-disease and population genetics, where the question is what a specific variant does rather than which target to pick.
A layered ablation of rule-based models for therapeutic exon-skipping targets in Duchenne muscular dystrophy. Starts from the reading-frame rule, then adds splice-topology and domain constraints and measures what each one buys. The published version corrects a substantive modelling error in the first draft — including one found by a reader whose child carries the deletion in question — and that correction is the scientific content.
shipped with its own correctionsWhere do disease mutations sit on the sodium–potassium pump, and can you tell without arguing in a circle? A reproducible integration of cross-species conservation, experimental structures, ClinVar curation, gnomAD constraint and AlphaMissense predictions across the four human Na+/K+-ATPase α-subunit paralogues.
Tools
falsifiable-targets-workflow (batch auditing at scale — Nextflow + Snakemake DAGs and a Docker image), ncbi-bioscraper (zero-cost PubMed + OpenAlex + open-access full-text mining), miniprotein_genai (generative miniprotein binder design in four Colab notebooks), and an MCAT concept-practice app for premeds.
// background
Education
Biomathematics, bioinformatics, and computational biology.
Genetics, genomics & cell biology; immunochemistry and cell culture. Thesis: Refining the epigenome through CRISPR-mediated techniques to establish a programmable system for transcriptional memory.
Research & lab
-
2018–21
Pines Lab Undergraduate Researcher — UC Berkeley
Magnetic-resonance research in the Pines Lab: zero- to ultra-low-field (ZULF) NMR and imaging methods that remove the need for strong magnetic fields. Worked with laser-polarized xenon molecular sensors, solid-state NMR of NV-diamond materials, optical hyperpolarization for signal enhancement, and miniaturized NMR detectors for portable bioimaging.
-
2021–22
Medical Technician — Discover Labs
Real-time PCR panels and automated RNA extraction under CLIA/HIPAA, with antimicrobial-treatment guidance and QC-workflow improvements.
-
2022
Microbiology Technician — Varian
Assessed non-invasive cancer therapies and ran viable / non-viable cleanroom monitoring under GLP, with supporting data analysis.
-
2018
Research Assistant — Tecnológico de Monterrey (Centro del Agua)
Microalgae bioprocessing at the Centro del Agua water-research center: spirulina cultivation, phycocyanin and DHA microencapsulation, and a microalgae-based UV-protective cream.
Industry & commercial
-
2024–25
Inside Sales Specialist — EditCo Bio (CRISPR)
Sole inside rep for a CRISPR reagents company across 30 states — a bench-trained scientist supporting researchers through their experiments, solving protocol problems rather than only closing deals.
-
2022–23
Sales Account Manager — GenScript
Managed a synthetic-biology product portfolio and coordinated custom gene-synthesis projects across pharma, biotech, and academic accounts.
-
2022
Consultant — PSC Biotech (Moderna)
IQ/OQ equipment qualification, sterilizer SOPs, and cleanroom quality control for Moderna manufacturing equipment under FDA and SAP guidelines.
-
2020–22
Founder & CEO — Creative Science
Founded and ran a cross-border resale business for lab and medical equipment (US / Mexico), sourced from pharma and biotech auctions.
// how i work
Method
- ▸Pre-registered, falsifiable hypotheses — fixed before the analysis runs.
- ▸Content-addressed, SHA-locked verdicts, so a result can't silently change.
- ▸Refusal-first — the pipeline is allowed to say "not enough evidence," and gaps are documented rather than smoothed over.
- ▸Null results shipped next to the positive ones.
- ▸Reproducible on free compute (Kaggle T4), in self-contained notebooks.