We compete with closed proprietary platforms on transparency. This page is where a sceptical reader can verify Liganx's scientific claims end-to-end — the 42-target Astex pose-accuracy benchmark and the per-pose confidence badge below, the docking-pocket coordinates we use, and the eleven literature-anchored mutation/drug pairs we run as positive controls. Every number below is regenerable from the linked code, and the cases where our pipeline can't resolve a published direction are surfaced explicitly with the structural-biology reason — not buried, not omitted.
Snapshot refreshed 131 days ago (2026-05-12 08:42 UTC) · Verification scripts available on request — write to hello@liganx.com
The number the field reports: for each target, re-dock the crystal ligand and ask whether the top-ranked pose lands within 2 Å of the experimental one. We run 42 targets of the Astex Diverse Set through one frozen protocol — docking box from the crystal ligand, exhaustiveness 16, single chain, no manual curation — the same fully automated path a user gets. (The full Astex Diverse Set is 85 complexes; we report the 42 our automated pipeline fetches and prepares end-to-end without hand-editing.) Every number below was computed live with the result cache off.
The gain comes from building the receptor from the crystal's declared biological assembly and keeping the metal ions in the pocket, instead of docking one bare chain. It changes the receptor on 17 of the 42 targets and is identical on the other 25. On those 17, over two independent replicates, hits went from 14/34 to 24/34. Four dimer-interface pockets flip from clearly wrong to clearly right in both replicates — 1HWI, 1SG0 and 1T9B (~6–7 Å → ≤ 1.5 Å) and 1Q1G — and 1KZK, an HIV-protease dimer, drops from 5.4 Å to 2.3/1.6 Å. Nothing regresses reproducibly. Assembly prep is now the default for Vina and GNINA.
Every per-target score, both preparations: download the 42-target CSV and recompute the 69% yourself.
How to read this against the literature: published redocking numbers — GNINA 1.0's own paper (~73%, PDBbind core), GNINA in PoseBusters v2 (~63%), AutoDock Vina (~44–58%) — use different benchmark sets with hand-curated inputs, so they are a directional band, not a head-to-head. Astex is an older, generally easier set. The honest, apples-to-apples comparison is our own automated pipeline against itself: 60 → 69 with the preparation fix.
Most co-folding tools report one average accuracy. Almost none tell you, for the pose in front of you, whether to believe it. Boltz-2 emits a per-ligand confidence (pLDDT); we bin it green / amber / red. On a 30-target holdout the thresholds were never fit to — overall base rate 15/30 (50%) correct, point-biserial r = 0.66 between confidence and correctness — the badge separates right from wrong cleanly:
| Badge | 30-target holdout | Independent 12-target run | What to do |
|---|---|---|---|
| Green · pLDDT ≥ 90 | 12 / 13 right (92%) | 5 / 5 right, all < 0.5 Å | Trust it |
| Amber · 70–90 | 3 / 8 right (38%) | 1 / 2 right | Look, don't bet |
| Red · < 70 | 0 / 9 right (0%) | 1 / 5 right | Don't use the pose |
Boltz-2 is run single-sequence (empty MSA, so wild type and mutants are predicted under identical conditions), with pocket conditioning on the residues inside the crystal-ligand box, one diffusion sample. The independent 12-target run reproduced the shape. One red pose (pLDDT 35) did land at 1.8 Å — which is why red means "don't rely on it," not "always wrong." This per-pose signal, not the raw success rate, is what makes an automated docking result actionable.
One rung above docking: an MM/GBSA binding-energy estimate on a reference pose, correlated against experimental affinity across eight congeneric series (184 ligands). The score is a single snapshot; a separate short MD run is a pose-stability trust badge, not averaged into the number. Read it as a ranking signal for ordering compounds within a target — not a calibrated Kd.
Rank correlation of the MM/GBSA score against experimental affinity, per target. We publish every target, including the weak ones (HIF2A, Thrombin, TNKS2).
| Target | Ligands | Pearson r | Spearman ρ |
|---|---|---|---|
| TYK2 | 13 | 0.82 | 0.83 |
| CDK2 | 10 | 0.72 | 0.79 |
| MCL1 | 25 | 0.76 | 0.77 |
| PTP1B | 22 | 0.67 | 0.75 |
| p38 | 29 | 0.48 | 0.67 |
| Thrombin | 21 | 0.65 | 0.54 |
| TNKS2 | 27 | 0.66 | 0.53 |
| HIF2A | 37 | 0.39 | 0.46 |
| Median · 8 targets | 184 | 0.67 | 0.71 |
We lead with rank correlation — the honest measure of a rescoring signal — and report every target, including the weak ones (HIF2A ρ = 0.46, Thrombin ρ = 0.54). A per-target calibrated RMSE (≈0.8 kcal/mol) exists but is deliberately kept out of the headline: it linearly maps scores onto the experimental scale, so it measures within-target ranking, not absolute affinity, and is easily misread as a Kd. Solute dielectric ε = 1 (tuned for ranking, not a physical constant); no explicit salt. FEP (relative free-energy perturbation) is in the pipeline for the close calls; we are not reporting FEP accuracy until a rigorous apples-to-apples benchmark is complete.
Every catalog target's pocket centre is independently verified against the chain-A co-crystal ligand centroid in the canonical RCSB PDB. Every promoted mutation is checked for residue identity and reachability inside the docking box. The check runs as a blocking gate before every backend deploy — a regression cannot ship.
Verification script: backend/scripts/verify_catalog.py · runs in CI before every deploy
A complementary, drug-focused check alongside the 42-target Astex run above: pull each drug out of its own co-crystal structure, re-dock it into the same pocket, and measure the RMSD to the experimental pose. Under 2 Å is the field's accepted pass mark. This is redocking (self-docking), the standard first check — not blind cross-docking: the search box is the target's pre-defined catalog box or, where none exists, a 22 Å cube auto-centred on the co-crystal ligand, so the pocket location is given. Exhaustiveness 16, one run per cell. Eight crystallographic complexes spanning ABL, KIT, EGFR, BRAF, ERα and trypsin — every number below was computed live with the result cache switched off, so nothing is recycled from an earlier run.
| Target | Drug | Vina (Å) | GNINA (Å) | Boltz-2 (Å) | Verdict |
|---|---|---|---|---|---|
| BRAF (3OG7) | Vemurafenib | 0.48 | 0.47 | 0.41 | PASS |
| KIT (1T46) | Imatinib | 1.01 | 0.99 | 0.35 | PASS |
| Trypsin (3PTB) | Benzamidine | 0.41 | 0.41 | 0.51 | PASS |
| ERα (3ERT) | 4-OH-Tamoxifen | 1.30 | 1.30 | 1.08 | PASS |
| ABL (2HYY) | Imatinib | 1.23 | 0.91 | 30.4 | PASS |
| EGFR (2ITY) | Gefitinib | 6.62 | 1.66 | 1.41 | PASS |
| EGFR (1M17) | Erlotinib | 7.75 | 1.97 | 1.50 | PASS |
| EGFR (1XKK) | Lapatinib | 2.91 | 3.07 | 2.93 | NOISE |
The honest read: fast Vina reproduces the crystal pose to well under 2 Å on 5 of 8 and misplaces the ligand on two EGFR ATP-site boxes (erlotinib 7.8 Å, gefitinib 6.6 Å). GNINA's CNN rescoring pulls both of those back into the pocket (→ 2.0 and 1.7 Å) for 7/8 under 2 Å. Boltz-2 co-folding, scored after pocket superposition, is the strongest on the well-behaved cases (KIT/imatinib 0.35 Å) at 6/8 under 2 Å — its one real miss is ABL/imatinib (30 Å), where it appears to fold ABL in the DFG-in state that imatinib can't occupy (the identical imatinib on KIT lands at 0.35 Å). Lapatinib is the lone case all three read as borderline (~3 Å). Verdict is PASS when any engine places the pose under 2 Å; NOISE when none clear 2.
Scripts: rescore.py (Vina / GNINA) · boltz_aligned.py (Boltz-2 pocket-aligned) · symmetry-corrected RMSD via spyRMSD · cache off, fresh GPU computes 28 Aug 2026
Time to a finished, interaction-annotated pose for one compound on a warm GPU worker (single-compound, no queue). Physics docking returns in well under a minute; Boltz-2 folds the entire protein, so it costs minutes — the price of a co-folded structure. Jobs fan across multiple GPU workers, so a whole screen finishes far faster than these per-pose numbers imply.
| Engine / method | Per-compound | Notes |
|---|---|---|
| Liganx — QuickVina2-GPU | ~35 s | our physics baseline; pure GPU search is a few seconds, rest is prep + interaction analysis |
| Liganx — GNINA (CNN) | ~34 s | our accuracy default; adds CNN pose rescoring |
| Liganx — Boltz-2 | ~2.2 min | our co-folding engine; whole-complex structure prediction |
| AutoDock Vina (CPU) | 1–15 min | the classic baseline on a CPU core |
| Uni-Dock / Vina-GPU | sub-sec–few s | GPU Vina forks; ~10³× CPU Vina in large batches |
| Schrödinger Glide SP | seconds–min | licensed physics docking, per core |
| DiffDock (ML) | ~10–40 s | diffusion pose generator, per complex on GPU |
| Boltz-2 / Chai (co-fold) | ~1–5 min | co-folding models; same class as our Boltz-2 |
Where we sit: our GPU Vina/GNINA are in the seconds-per-compound band with Uni-Dock and Vina-GPU and comfortably ahead of CPU Vina and licensed Glide, while Boltz-2 is in the same minutes-per-complex band as every other co-folding model. The advantage is offering all three — fast screen, CNN-rescored accuracy, and full co-folding — behind one API with no license and automatic multi-GPU fan-out. Competitor figures are indicative ranges from published tooling, not head-to-head runs on identical hardware.
Three more known-answer checks on the surrounding engines, reported at face value — modest where the method is modest.
Screening enrichment is honest but small-sample (24 compounds on a single ABL target) — read it as a sanity check that actives rank above decoys, not as a DUD-E-scale claim. ADMET is rule/ML-based triage scored against public benchmarks: useful for flagging obvious liabilities (high sensitivity on hERG and BBB), weak on the harder directional splits. Throughput is the reliability work from this sprint — a job now fans its compounds across multiple GPU workers concurrently, and the per-job GPU-worker count is shown on every result and in history.
Scripts: vs_analyze.py · admet/eval.py · benchmarked Aug 2026 on production liganx-api
The summary above is honest but compressed. Here's the longer version a chemist would want before trusting a Δ value from this pipeline on their own mutation/drug pair.
On 2026-05-01 a PhD-level audit suggested our wild-type and mutant receptors should go through the same OpenMM amber99sb-ildn vacuum minimisation. We implemented the symmetric prep, re-ran this suite, and watched 5/8 PASS collapse to 2/8 PASS + 1 FAIL. We reverted the change the same day. The lesson: WT comes from a crystal structure (already a low-energy minimum) while mutant comes from a synthetic side-chain swap that needs relaxation — the asymmetry is correct, not a bug. We've added a CI gate that fails the build if anyone re-introduces the symmetric-prep pattern. Validation suites only work if they're allowed to overrule a smart-sounding prior; ours did, and we acted on it.
The validation suite below is the data backing every claim in the two columns above. Each row links to a live job you can re-open and inspect — pose, contacts, 2D map — and a verdict_note that explains the case-specific reasoning. Cases with documented method limits carry a caveat field that surfaces the structural biology behind why a particular Δ does or doesn't match its literature direction.
Eleven (target, mutation, drug) pairs whose mutation-driven binding shifts are published in the clinical and pharmacology literature. We submit each to the live Liganx pipeline at exhaustiveness=16 (2× the product default — tighter sampling for a tighter noise band), capture Δ(mutant − WT), and check whether the direction agrees with the published shift. The point is direction, not magnitude — Vina scoring isn't free energy and isn't calibrated to cellular IC50.
Noise floor: ±1 kcal/mol (Vina scoring reproducibility at default exhaustiveness). NOISE means the Δ direction may be correct but magnitude sits below the noise floor — see the per-case rationale. FAIL means the Δ direction explicitly disagrees with the published literature direction; that's a regression and we investigate.
| Case | Drug | WT | Mut | Δ | Expected | Verdict |
|---|---|---|---|---|---|---|
| ABL T315I — Imatinib resistance | Imatinib | -10.00 | -8.80 | +1.20 | resistance | PASS |
| EGFR T790M — Gefitinib resistance | Gefitinib | -7.70 | -6.40 | +1.30 | resistance | PASS |
| EGFR T790M — Osimertinib retains/gains | Osimertinib | -7.20 | -7.00 | +0.20 | selectivity | NOISE |
| BRAF V600E — Vemurafenib selectivity | Vemurafenib | -8.70 | -8.70 | +0.00 | selectivity | NOISE |
| KIT D816V — Imatinib resistance | Imatinib | -13.10 | -10.10 | +3.00 | resistance | PASS |
| KIT D816V — Avapritinib selectivity | Avapritinib | -8.50 | -8.90 | -0.40 | selectivity | NOISE |
| BTK C481S — Ibrutinib covalent escape | Ibrutinib | -10.00 | -8.20 | +1.80 | resistance | PASS |
| BTK C481S — Pirtobrutinib retention | Pirtobrutinib | -9.90 | -9.40 | +0.50 | retained | PASS |
| KRAS G12C — Sotorasib selectivity | Sotorasib | -7.00 | -7.00 | +0.00 | selectivity | NOISE |
| EGFR C797S — Osimertinib resistance | Osimertinib | -8.00 | -7.40 | +0.60 | resistance | NOISE |
| EGFR L858R — Gefitinib selectivity | Gefitinib | -8.00 | -6.80 | +1.20 | selectivity | FAIL |
Validation script: backend/scripts/validate_positive_controls.py · each row's case name links to the live Liganx job result on this deployment.
One paragraph per case explaining the published literature, the number Liganx returned, and why a NOISE result (where applicable) is a method limitation rather than a bug.
Literature: O'Hare et al., Nat Rev Cancer 2007 — T315I drives near-total Imatinib loss in CML; >400-fold IC50 shift in cellular assays.
Result: Δ=+1.20 kcal/mol matches expected 'resistance'
Literature: Pao et al., PLoS Med 2005 — T790M gatekeeper drives 1st-gen TKI resistance in NSCLC.
Result: Δ=+1.30 kcal/mol matches expected 'resistance'
Literature: Cross et al., Cancer Discov 2014 — Osimertinib (AZD9291) was specifically designed to retain potency against T790M; cellular IC50 ~1 nM vs T790M, comparable to or better than WT.
Caveat: Osimertinib is a covalent acrylamide on C797. Vina is non-covalent — magnitude will be smaller than IC50 shifts suggest, but the geometric T790M preference (the methionine bulk fits the Osimertinib scaffold better than threonine does) still produces a measurable ΔΔG.
Result: |Δ|=0.20 below ±1.0 kcal/mol noise floor — direction not resolvable at exhaustiveness=16
Literature: Bollag et al., Nature 2010 — Vemurafenib (PLX4032) is V600E-selective; ~30 nM against V600E vs ~100 nM against WT BRAF in cellular assays.
Result: |Δ|=0.00 below ±1.0 kcal/mol noise floor — direction not resolvable at exhaustiveness=16
Literature: Heinrich et al., J Clin Oncol 2003 — D816V shifts KIT into an active conformation Imatinib cannot bind; >100-fold loss in mastocytosis.
Result: Δ=+3.00 kcal/mol matches expected 'resistance'
Literature: Evans et al., Sci Transl Med 2017 — Avapritinib (BLU-285) was designed to bind the D816V active conformation; sub-nM cellular potency.
Result: |Δ|=0.40 below ±1.0 kcal/mol noise floor — direction not resolvable at exhaustiveness=16
Literature: Woyach et al., NEJM 2014 — C481S ablates the covalent cysteine target of Ibrutinib; >100-fold cellular IC50 loss in CLL.
Caveat: Ibrutinib's WT advantage is COVALENT (acrylamide → C481 thiol). Vina is non-covalent and will under-represent this effect; the residual non-covalent ΔΔG should still be positive (S is smaller than C, slight pocket reshape) but smaller than IC50 data suggests.
Result: Δ=+1.80 kcal/mol matches expected 'resistance'
Literature: Mato et al., NEJM 2021 — Pirtobrutinib is non-covalent; retains potency against C481S (cellular IC50 essentially unchanged).
Caveat: Pirtobrutinib's clinical retention against C481S comes from being non-covalent — no acrylamide to lose. Rigid-receptor Vina cannot model covalent vs non-covalent mechanism, so Pirtobrutinib and Ibrutinib both return similar small-magnitude ΔΔG against C481S even though their clinical IC50 shifts differ by >100×. Documented limitation, not a pipeline regression.
Result: |Δ|=0.50 within ±1.0 kcal/mol noise floor (retained as expected)
Literature: Canon et al., Nature 2019 — Sotorasib (AMG 510) was specifically designed to bind GDP-bound KRAS G12C; cellular IC50 ~5-15 nM against G12C and effectively inert against WT. First FDA-approved direct KRAS inhibitor (May 2021).
Caveat: Sotorasib is a COVALENT acrylamide targeting the engineered C12 thiol via Michael addition. Vina is rigid-receptor non-covalent; it cannot model the covalent attack or the cryptic switch-II conformational opening.
Result: |Δ|=0.00 below ±1.0 kcal/mol noise floor
Literature: Thress et al., Nat Med 2015 — C797S is the on-target acquired resistance mutation to Osimertinib (3rd-gen EGFR TKI). Loss of the covalent cysteine ablates the irreversible bond; cellular IC50 shifts >100×.
Caveat: Osimertinib's WT advantage is COVALENT (acrylamide → C797 thiol). C797S is precisely the loss-of-target Vina cannot see; both WT and mutant score similarly under non-covalent scoring.
Result: |Δ|=0.60 below ±1.0 kcal/mol noise floor
Literature: Lynch et al., NEJM 2004 — L858R activating mutation drives 1st-gen TKI response in NSCLC. Gefitinib has ~10× higher affinity for the L858R-activated kinase than for WT EGFR.
Caveat: L858R activates the kinase via conformational shift. Rigid-receptor Vina docks against the static crystal pocket and under-represents state-equilibrium effects.
Result: Δ=+1.20 disagrees with 'selectivity' — documented method limitation.
Five of the eleven cases above return a Δ smaller than ±1 kcal/mol — below the reproducibility floor of Vina scoring at default exhaustiveness. The direction is usually correct but the magnitude is sub-noise, so we report them honestly as NOISE rather than claiming a confident answer. A sixth case (EGFR L858R + Gefitinib) returns the wrong direction at above-noise magnitude — a true FAIL we surface here because L858R's selectivity is conformational and rigid-receptor docking against a static crystal pocket cannot capture state-equilibrium effects. Documented method limit, not a pipeline regression.
The dominant cause is that our default mutant-receptor builder applies the residue substitution but does not energy- minimise the surrounding side chains. WT and mutant receptors therefore differ only at one side chain, with no global pocket reshape. For mutations whose biological effect is conformational (KRAS Q61, KIT D816V, BRAF V600 in the inactive state), the rigid-receptor docking simply cannot resolve the shift — even though the catalog gets the docking-pocket coordinates exactly right and the pipeline runs end-to-end without errors.
FoldX BuildModel does minimise the structure and produces a measurable Δ for these cases. We ship the FoldX call path in the runner; the production image runs PDBFixer-only because FoldX's academic-only licence isn't compatible with a public web service. Restoring FoldX behind an opt-in for academic users is on the roadmap.
Two cases with sub-noise Vina scores (BTK / Ibrutinib and EGFR / Osimertinib) are known to be covalent-inhibitor cases: their WT vs mutant advantage is partly the covalent bond to a cysteine, which a non-covalent docking engine cannot model. Vina is not the right tool for those Δs — the result is not "Liganx is wrong" but "Liganx + Vina is the wrong tool here, use free-energy methods (FEP) or a covalent docker like CovDock".
We publish these limitations on purpose. A platform that pretends to give a confident answer where the underlying method can't is worse than one that says "below noise — not interpretable" and tells you which kinds of mutations Vina is and isn't the right tool for.