Scientific validation

Show your work.

We compete with closed proprietary platforms on transparency. This page is where a sceptical reader can verify Liganx's scientific claims end-to-end — the 42-target Astex pose-accuracy benchmark and the per-pose confidence badge below, the docking-pocket coordinates we use, and the eleven literature-anchored mutation/drug pairs we run as positive controls. Every number below is regenerable from the linked code, and the cases where our pipeline can't resolve a published direction are surfaced explicitly with the structural-biology reason — not buried, not omitted.

Snapshot refreshed 131 days ago (2026-05-12 08:42 UTC) · Verification scripts available on request — write to hello@liganx.com

Pose accuracy — 42-target Astex benchmark

The number the field reports: for each target, re-dock the crystal ligand and ask whether the top-ranked pose lands within 2 Å of the experimental one. We run 42 targets of the Astex Diverse Set through one frozen protocol — docking box from the crystal ligand, exhaustiveness 16, single chain, no manual curation — the same fully automated path a user gets. (The full Astex Diverse Set is 85 complexes; we report the 42 our automated pipeline fetches and prepares end-to-end without hand-editing.) Every number below was computed live with the result cache off.

GNINA · assembly prep
29 / 42
69% top-1 < 2 Å · biological assembly + pocket metals (now the default)
GNINA · previous prep
25 / 42
60% · same-day paired run · single chain, no metals
GNINA · reproducibility
63% ± 0.9
26/27/26 of 42 across three seeds (Aug 29)
QuickVina2-GPU
32%
fast physics baseline, no CNN rescoring · Boltz-2 co-folding 35–38%
What moved the number: preparation, not the engine

The gain comes from building the receptor from the crystal's declared biological assembly and keeping the metal ions in the pocket, instead of docking one bare chain. It changes the receptor on 17 of the 42 targets and is identical on the other 25. On those 17, over two independent replicates, hits went from 14/34 to 24/34. Four dimer-interface pockets flip from clearly wrong to clearly right in both replicates — 1HWI, 1SG0 and 1T9B (~6–7 Å → ≤ 1.5 Å) and 1Q1G — and 1KZK, an HIV-protease dimer, drops from 5.4 Å to 2.3/1.6 Å. Nothing regresses reproducibly. Assembly prep is now the default for Vina and GNINA.

Every per-target score, both preparations: download the 42-target CSV and recompute the 69% yourself.

How to read this against the literature: published redocking numbers — GNINA 1.0's own paper (~73%, PDBbind core), GNINA in PoseBusters v2 (~63%), AutoDock Vina (~44–58%) — use different benchmark sets with hand-curated inputs, so they are a directional band, not a head-to-head. Astex is an older, generally easier set. The honest, apples-to-apples comparison is our own automated pipeline against itself: 60 → 69 with the preparation fix.

The confidence badge — which pose to trust

Most co-folding tools report one average accuracy. Almost none tell you, for the pose in front of you, whether to believe it. Boltz-2 emits a per-ligand confidence (pLDDT); we bin it green / amber / red. On a 30-target holdout the thresholds were never fit to — overall base rate 15/30 (50%) correct, point-biserial r = 0.66 between confidence and correctness — the badge separates right from wrong cleanly:

Badge30-target holdoutIndependent 12-target runWhat to do
Green · pLDDT ≥ 9012 / 13 right (92%)5 / 5 right, all < 0.5 ÅTrust it
Amber · 70–903 / 8 right (38%)1 / 2 rightLook, don't bet
Red · < 700 / 9 right (0%)1 / 5 rightDon't use the pose

Boltz-2 is run single-sequence (empty MSA, so wild type and mutants are predicted under identical conditions), with pocket conditioning on the residues inside the crystal-ligand box, one diffusion sample. The independent 12-target run reproduced the shape. One red pose (pLDDT 35) did land at 1.8 Å — which is why red means "don't rely on it," not "always wrong." This per-pose signal, not the raw success rate, is what makes an automated docking result actionable.

Affinity ranking — 8-target MM/GBSA benchmark

One rung above docking: an MM/GBSA binding-energy estimate on a reference pose, correlated against experimental affinity across eight congeneric series (184 ligands). The score is a single snapshot; a separate short MD run is a pose-stability trust badge, not averaged into the number. Read it as a ranking signal for ordering compounds within a target — not a calibrated Kd.

MM/GBSA rank correlation8 targets · 184 ligandsmedian Spearman ρ 0.71
00.51.0Spearman ρ — rank correlation vs experiment (1.0 = perfect)TYK20.83CDK20.79MCL10.77PTP1B0.75p380.67Thrombin0.54TNKS20.53HIF2A0.46median 0.71

Rank correlation of the MM/GBSA score against experimental affinity, per target. We publish every target, including the weak ones (HIF2A, Thrombin, TNKS2).

TargetLigandsPearson rSpearman ρ
TYK2130.820.83
CDK2100.720.79
MCL1250.760.77
PTP1B220.670.75
p38290.480.67
Thrombin210.650.54
TNKS2270.660.53
HIF2A370.390.46
Median · 8 targets1840.670.71

We lead with rank correlation — the honest measure of a rescoring signal — and report every target, including the weak ones (HIF2A ρ = 0.46, Thrombin ρ = 0.54). A per-target calibrated RMSE (≈0.8 kcal/mol) exists but is deliberately kept out of the headline: it linearly maps scores onto the experimental scale, so it measures within-target ranking, not absolute affinity, and is easily misread as a Kd. Solute dielectric ε = 1 (tuned for ranking, not a physical constant); no explicit salt. FEP (relative free-energy perturbation) is in the pipeline for the close calls; we are not reporting FEP accuracy until a rigorous apples-to-apples benchmark is complete.

Catalog audit

Every catalog target's pocket centre is independently verified against the chain-A co-crystal ligand centroid in the canonical RCSB PDB. Every promoted mutation is checked for residue identity and reachability inside the docking box. The check runs as a blocking gate before every backend deploy — a regression cannot ship.

Catalog targets
13
all centres ≤ 5 Å from canonical ligand
Promoted mutations
40
across 13 oncology targets
Reachable in pocket
32 / 40
8 documented out-of-scope (allosteric, helical-domain, switch-region)

Verification script: backend/scripts/verify_catalog.py · runs in CI before every deploy

Approved-drug redocking — 8 complexes

A complementary, drug-focused check alongside the 42-target Astex run above: pull each drug out of its own co-crystal structure, re-dock it into the same pocket, and measure the RMSD to the experimental pose. Under 2 Å is the field's accepted pass mark. This is redocking (self-docking), the standard first check — not blind cross-docking: the search box is the target's pre-defined catalog box or, where none exists, a 22 Å cube auto-centred on the co-crystal ligand, so the pocket location is given. Exhaustiveness 16, one run per cell. Eight crystallographic complexes spanning ABL, KIT, EGFR, BRAF, ERα and trypsin — every number below was computed live with the result cache switched off, so nothing is recycled from an earlier run.

Vina · QuickVina2-GPU
5 / 8 < 2 Å
median 1.27 Å · fast physics baseline · misses two ATP-site boxes
GNINA · CNN rescoring
7 / 8 < 2 Å
median 1.15 Å · rescues both Vina misses (7.8→2.0, 6.6→1.7 Å)
Boltz-2 · co-folding
6 / 8 < 2 Å
median 1.25 Å pocket-aligned · 1 outlier (ABL/imatinib)
TargetDrugVina (Å)GNINA (Å)Boltz-2 (Å)Verdict
BRAF (3OG7)Vemurafenib0.480.470.41PASS
KIT (1T46)Imatinib1.010.990.35PASS
Trypsin (3PTB)Benzamidine0.410.410.51PASS
ERα (3ERT)4-OH-Tamoxifen1.301.301.08PASS
ABL (2HYY)Imatinib1.230.9130.4PASS
EGFR (2ITY)Gefitinib6.621.661.41PASS
EGFR (1M17)Erlotinib7.751.971.50PASS
EGFR (1XKK)Lapatinib2.913.072.93NOISE

The honest read: fast Vina reproduces the crystal pose to well under 2 Å on 5 of 8 and misplaces the ligand on two EGFR ATP-site boxes (erlotinib 7.8 Å, gefitinib 6.6 Å). GNINA's CNN rescoring pulls both of those back into the pocket (→ 2.0 and 1.7 Å) for 7/8 under 2 Å. Boltz-2 co-folding, scored after pocket superposition, is the strongest on the well-behaved cases (KIT/imatinib 0.35 Å) at 6/8 under 2 Å — its one real miss is ABL/imatinib (30 Å), where it appears to fold ABL in the DFG-in state that imatinib can't occupy (the identical imatinib on KIT lands at 0.35 Å). Lapatinib is the lone case all three read as borderline (~3 Å). Verdict is PASS when any engine places the pose under 2 Å; NOISE when none clear 2.

Scripts: rescore.py (Vina / GNINA) · boltz_aligned.py (Boltz-2 pocket-aligned) · symmetry-corrected RMSD via spyRMSD · cache off, fresh GPU computes 28 Aug 2026

Docking speed

Time to a finished, interaction-annotated pose for one compound on a warm GPU worker (single-compound, no queue). Physics docking returns in well under a minute; Boltz-2 folds the entire protein, so it costs minutes — the price of a co-folded structure. Jobs fan across multiple GPU workers, so a whole screen finishes far faster than these per-pose numbers imply.

Vina (QuickVina2-GPU)
~35 s
prep + GPU dock + ProLIF contacts, per compound (warm)
GNINA (CNN)
~34 s
Vina search + CNN pose rescoring, per compound (warm)
Boltz-2 co-folding
~2.2 min
folds the full protein–ligand complex, per compound
Engine / methodPer-compoundNotes
Liganx — QuickVina2-GPU~35 sour physics baseline; pure GPU search is a few seconds, rest is prep + interaction analysis
Liganx — GNINA (CNN)~34 sour accuracy default; adds CNN pose rescoring
Liganx — Boltz-2~2.2 minour co-folding engine; whole-complex structure prediction
AutoDock Vina (CPU)1–15 minthe classic baseline on a CPU core
Uni-Dock / Vina-GPUsub-sec–few sGPU Vina forks; ~10³× CPU Vina in large batches
Schrödinger Glide SPseconds–minlicensed physics docking, per core
DiffDock (ML)~10–40 sdiffusion pose generator, per complex on GPU
Boltz-2 / Chai (co-fold)~1–5 minco-folding models; same class as our Boltz-2

Where we sit: our GPU Vina/GNINA are in the seconds-per-compound band with Uni-Dock and Vina-GPU and comfortably ahead of CPU Vina and licensed Glide, while Boltz-2 is in the same minutes-per-complex band as every other co-folding model. The advantage is offering all three — fast screen, CNN-rescored accuracy, and full co-folding — behind one API with no license and automatic multi-GPU fan-out. Competitor figures are indicative ranges from published tooling, not head-to-head runs on identical hardware.

Screening, ADMET & throughput

Three more known-answer checks on the surrounding engines, reported at face value — modest where the method is modest.

Virtual screening (ABL)
AUC 0.81
12 literature actives vs 12 property-matched decoys; actives fill 7 of the top 8 ranks
ADMET vs TDC
AUC 0.55–0.74
5 Therapeutics Data Commons sets (BBB, hERG, CYP3A4/2D6, DILI); directional flags, not calibrated probabilities
Parallel throughput
5 / 3 / 2
Vina bursts to 5 GPU workers, GNINA to 3, Boltz-2 to 2; 0 failed dispatches after the thundering-herd fix

Screening enrichment is honest but small-sample (24 compounds on a single ABL target) — read it as a sanity check that actives rank above decoys, not as a DUD-E-scale claim. ADMET is rule/ML-based triage scored against public benchmarks: useful for flagging obvious liabilities (high sensitivity on hERG and BBB), weak on the harder directional splits. Throughput is the reliability work from this sprint — a job now fans its compounds across multiple GPU workers concurrently, and the per-job GPU-worker count is shown on every result and in history.

Scripts: vs_analyze.py · admet/eval.py · benchmarked Aug 2026 on production liganx-api

What this means for your work

The summary above is honest but compressed. Here's the longer version a chemist would want before trusting a Δ value from this pipeline on their own mutation/drug pair.

When to trust this pipeline
  • Gatekeeper-residue resistance. ABL T315I + Imatinib (Δ +1.2 kcal/mol) and EGFR T790M + Gefitinib (Δ +1.3 kcal/mol) — both clean PASS. Steric clash from a single residue substitution is exactly what rigid-receptor Vina was designed to capture.
  • Active-conformation mutations against an inactive- conformation drug. KIT D816V + Imatinib — clean PASS at +3.0 kcal/mol, the strongest signal in the suite. The opposite of selectivity (Avapritinib, see right) and the case Vina is best suited for.
  • The non-covalent component of covalent escape.BTK C481S + Ibrutinib — PASS at +1.8 kcal/mol on the residual non-covalent ΔΔG. The covalent contribution to Ibrutinib's clinical loss is not modelled (Vina is non-covalent), but the geometric C→S signal is real.
  • Non-covalent retention against covalent escape.BTK C481S + Pirtobrutinib — PASS at Δ +0.5 kcal/mol, well within the ±1.0 noise floor that's the right answer for "no clinical potency loss." Pirtobrutinib is non-covalent so C481S doesn't disrupt its binding mode; the small in-pocket Δ matches the clinical retention.
When to be cautious
  • Covalent inhibitors that depend on the warhead for selectivity. Osimertinib (acrylamide on C797) and the covalent leg of Ibrutinib's C481 binding — Vina is non-covalent and cannot resolve covalent vs non-covalent mechanism. These show as NOISE-magnitude or qualitative-only on the table below.
  • Activation-loop selectivity that needs induced fit.BRAF V600E + Vemurafenib lands in the noise floor in our current snapshot — the 4WO5 receptor is in DFG-out but the V→E substitution alone (no induced-fit relaxation) doesn't fully shape the αC-helix-out state Vemurafenib was optimised for. The literature direction is well-established (~3-fold cellular IC50 shift) but our Δ stays at the noise boundary.
  • Active-conformation selectivity drugs. Avapritinib was designed to bind the D816V-stabilised active conformation; our 1T46 receptor is not in that conformation and induced-fit relaxation is outside the scope of rigid docking. Documented method limit.
  • Switch-region or allosteric mutations. Any mutation whose binding effect propagates through long-range conformational change (αC-helix flips, P-loop dynamics) is outside what a rigid receptor can capture. Catalog audit marks these as out-of-scope before they hit the docking step.
  • Absolute affinity prediction. Vina is empirical scoring, not free-energy perturbation. Read Δ values as direction at above-noise magnitude, not as ΔΔG of binding. For affinity estimates use alchemical free-energy methods (FEP / TI) validated against assay data — docking doesn't do that.
How this snapshot is honest about its own failure modes

On 2026-05-01 a PhD-level audit suggested our wild-type and mutant receptors should go through the same OpenMM amber99sb-ildn vacuum minimisation. We implemented the symmetric prep, re-ran this suite, and watched 5/8 PASS collapse to 2/8 PASS + 1 FAIL. We reverted the change the same day. The lesson: WT comes from a crystal structure (already a low-energy minimum) while mutant comes from a synthetic side-chain swap that needs relaxation — the asymmetry is correct, not a bug. We've added a CI gate that fails the build if anyone re-introduces the symmetric-prep pattern. Validation suites only work if they're allowed to overrule a smart-sounding prior; ours did, and we acted on it.

The validation suite below is the data backing every claim in the two columns above. Each row links to a live job you can re-open and inspect — pose, contacts, 2D map — and a verdict_note that explains the case-specific reasoning. Cases with documented method limits carry a caveat field that surfaces the structural biology behind why a particular Δ does or doesn't match its literature direction.

Positive-control validation

Eleven (target, mutation, drug) pairs whose mutation-driven binding shifts are published in the clinical and pharmacology literature. We submit each to the live Liganx pipeline at exhaustiveness=16 (2× the product default — tighter sampling for a tighter noise band), capture Δ(mutant − WT), and check whether the direction agrees with the published shift. The point is direction, not magnitude — Vina scoring isn't free energy and isn't calibrated to cellular IC50.

Pass
5
Noise
5
Skip
0
Fail
1

Noise floor: ±1 kcal/mol (Vina scoring reproducibility at default exhaustiveness). NOISE means the Δ direction may be correct but magnitude sits below the noise floor — see the per-case rationale. FAIL means the Δ direction explicitly disagrees with the published literature direction; that's a regression and we investigate.

CaseDrugWTMutΔExpectedVerdict
ABL T315I — Imatinib resistanceImatinib-10.00-8.80+1.20resistancePASS
EGFR T790M — Gefitinib resistanceGefitinib-7.70-6.40+1.30resistancePASS
EGFR T790M — Osimertinib retains/gainsOsimertinib-7.20-7.00+0.20selectivityNOISE
BRAF V600E — Vemurafenib selectivityVemurafenib-8.70-8.70+0.00selectivityNOISE
KIT D816V — Imatinib resistanceImatinib-13.10-10.10+3.00resistancePASS
KIT D816V — Avapritinib selectivityAvapritinib-8.50-8.90-0.40selectivityNOISE
BTK C481S — Ibrutinib covalent escapeIbrutinib-10.00-8.20+1.80resistancePASS
BTK C481S — Pirtobrutinib retentionPirtobrutinib-9.90-9.40+0.50retainedPASS
KRAS G12C — Sotorasib selectivitySotorasib-7.00-7.00+0.00selectivityNOISE
EGFR C797S — Osimertinib resistanceOsimertinib-8.00-7.40+0.60resistanceNOISE
EGFR L858R — Gefitinib selectivityGefitinib-8.00-6.80+1.20selectivityFAIL

Validation script: backend/scripts/validate_positive_controls.py · each row's case name links to the live Liganx job result on this deployment.

Case-by-case rationale

One paragraph per case explaining the published literature, the number Liganx returned, and why a NOISE result (where applicable) is a method limitation rather than a bug.

2HYY/A · T315I · ImatinibPASS

Literature: O'Hare et al., Nat Rev Cancer 2007 — T315I drives near-total Imatinib loss in CML; >400-fold IC50 shift in cellular assays.

Result: Δ=+1.20 kcal/mol matches expected 'resistance'

2ITY/A · T790M · GefitinibPASS

Literature: Pao et al., PLoS Med 2005 — T790M gatekeeper drives 1st-gen TKI resistance in NSCLC.

Result: Δ=+1.30 kcal/mol matches expected 'resistance'

2ITY/A · T790M · OsimertinibNOISE

Literature: Cross et al., Cancer Discov 2014 — Osimertinib (AZD9291) was specifically designed to retain potency against T790M; cellular IC50 ~1 nM vs T790M, comparable to or better than WT.

Caveat: Osimertinib is a covalent acrylamide on C797. Vina is non-covalent — magnitude will be smaller than IC50 shifts suggest, but the geometric T790M preference (the methionine bulk fits the Osimertinib scaffold better than threonine does) still produces a measurable ΔΔG.

Result: |Δ|=0.20 below ±1.0 kcal/mol noise floor — direction not resolvable at exhaustiveness=16

4WO5/A · V600E · VemurafenibNOISE

Literature: Bollag et al., Nature 2010 — Vemurafenib (PLX4032) is V600E-selective; ~30 nM against V600E vs ~100 nM against WT BRAF in cellular assays.

Result: |Δ|=0.00 below ±1.0 kcal/mol noise floor — direction not resolvable at exhaustiveness=16

1T46/A · D816V · ImatinibPASS

Literature: Heinrich et al., J Clin Oncol 2003 — D816V shifts KIT into an active conformation Imatinib cannot bind; >100-fold loss in mastocytosis.

Result: Δ=+3.00 kcal/mol matches expected 'resistance'

1T46/A · D816V · AvapritinibNOISE

Literature: Evans et al., Sci Transl Med 2017 — Avapritinib (BLU-285) was designed to bind the D816V active conformation; sub-nM cellular potency.

Result: |Δ|=0.40 below ±1.0 kcal/mol noise floor — direction not resolvable at exhaustiveness=16

5P9J/A · C481S · IbrutinibPASS

Literature: Woyach et al., NEJM 2014 — C481S ablates the covalent cysteine target of Ibrutinib; >100-fold cellular IC50 loss in CLL.

Caveat: Ibrutinib's WT advantage is COVALENT (acrylamide → C481 thiol). Vina is non-covalent and will under-represent this effect; the residual non-covalent ΔΔG should still be positive (S is smaller than C, slight pocket reshape) but smaller than IC50 data suggests.

Result: Δ=+1.80 kcal/mol matches expected 'resistance'

5P9J/A · C481S · PirtobrutinibPASS

Literature: Mato et al., NEJM 2021 — Pirtobrutinib is non-covalent; retains potency against C481S (cellular IC50 essentially unchanged).

Caveat: Pirtobrutinib's clinical retention against C481S comes from being non-covalent — no acrylamide to lose. Rigid-receptor Vina cannot model covalent vs non-covalent mechanism, so Pirtobrutinib and Ibrutinib both return similar small-magnitude ΔΔG against C481S even though their clinical IC50 shifts differ by >100×. Documented limitation, not a pipeline regression.

Result: |Δ|=0.50 within ±1.0 kcal/mol noise floor (retained as expected)

4OBE/A · G12C · SotorasibNOISE

Literature: Canon et al., Nature 2019 — Sotorasib (AMG 510) was specifically designed to bind GDP-bound KRAS G12C; cellular IC50 ~5-15 nM against G12C and effectively inert against WT. First FDA-approved direct KRAS inhibitor (May 2021).

Caveat: Sotorasib is a COVALENT acrylamide targeting the engineered C12 thiol via Michael addition. Vina is rigid-receptor non-covalent; it cannot model the covalent attack or the cryptic switch-II conformational opening.

Result: |Δ|=0.00 below ±1.0 kcal/mol noise floor

2ITY/A · C797S · OsimertinibNOISE

Literature: Thress et al., Nat Med 2015 — C797S is the on-target acquired resistance mutation to Osimertinib (3rd-gen EGFR TKI). Loss of the covalent cysteine ablates the irreversible bond; cellular IC50 shifts >100×.

Caveat: Osimertinib's WT advantage is COVALENT (acrylamide → C797 thiol). C797S is precisely the loss-of-target Vina cannot see; both WT and mutant score similarly under non-covalent scoring.

Result: |Δ|=0.60 below ±1.0 kcal/mol noise floor

2ITY/A · L858R · GefitinibFAIL

Literature: Lynch et al., NEJM 2004 — L858R activating mutation drives 1st-gen TKI response in NSCLC. Gefitinib has ~10× higher affinity for the L858R-activated kinase than for WT EGFR.

Caveat: L858R activates the kinase via conformational shift. Rigid-receptor Vina docks against the static crystal pocket and under-represents state-equilibrium effects.

Result: Δ=+1.20 disagrees with 'selectivity' — documented method limitation.

Why several cases land below the noise floor

Five of the eleven cases above return a Δ smaller than ±1 kcal/mol — below the reproducibility floor of Vina scoring at default exhaustiveness. The direction is usually correct but the magnitude is sub-noise, so we report them honestly as NOISE rather than claiming a confident answer. A sixth case (EGFR L858R + Gefitinib) returns the wrong direction at above-noise magnitude — a true FAIL we surface here because L858R's selectivity is conformational and rigid-receptor docking against a static crystal pocket cannot capture state-equilibrium effects. Documented method limit, not a pipeline regression.

The dominant cause is that our default mutant-receptor builder applies the residue substitution but does not energy- minimise the surrounding side chains. WT and mutant receptors therefore differ only at one side chain, with no global pocket reshape. For mutations whose biological effect is conformational (KRAS Q61, KIT D816V, BRAF V600 in the inactive state), the rigid-receptor docking simply cannot resolve the shift — even though the catalog gets the docking-pocket coordinates exactly right and the pipeline runs end-to-end without errors.

FoldX BuildModel does minimise the structure and produces a measurable Δ for these cases. We ship the FoldX call path in the runner; the production image runs PDBFixer-only because FoldX's academic-only licence isn't compatible with a public web service. Restoring FoldX behind an opt-in for academic users is on the roadmap.

Two cases with sub-noise Vina scores (BTK / Ibrutinib and EGFR / Osimertinib) are known to be covalent-inhibitor cases: their WT vs mutant advantage is partly the covalent bond to a cysteine, which a non-covalent docking engine cannot model. Vina is not the right tool for those Δs — the result is not "Liganx is wrong" but "Liganx + Vina is the wrong tool here, use free-energy methods (FEP) or a covalent docker like CovDock".

We publish these limitations on purpose. A platform that pretends to give a confident answer where the underlying method can't is worse than one that says "below noise — not interpretable" and tells you which kinds of mutations Vina is and isn't the right tool for.

Run a docking job