ACROBiosystems · AIDD 抗体验证服务ACROBiosystems · AIDD Antibody Validation
AI 一天能设计上千条序列,湿实验一轮验证要 7–10 天。CHODirect™ 把这一轮压到快至 4 天——真实全长 IgG、CHO 细胞表达、上清直接上机不纯化,交付的是能直接进训练管线的结构化数据。
AI designs a thousand sequences a day; the wet lab needs 7–10 days per validation round. CHODirect™ compresses that round to as few as 4 days — real full-length IgG, expressed in CHO, supernatant straight onto the instrument with no purification. What comes back is structured data that goes directly into your training pipeline.
01 问题Problem
模型迭代速度已经远超湿实验。验证环节每慢一天,模型就少拿一天的真实反馈。Model iteration has long outpaced the wet lab. Every day validation lags is a day the model goes without real feedback.
传统 CHO 表达加纯化要 7 到 10 天。模型一天产出成百上千条序列,实验验证成了唯一的节流阀。
Conventional CHO expression plus purification takes 7 to 10 days. A model produces hundreds of sequences a day, so experimental validation becomes the only throttle in the loop.
低活性、错误构象的靶点蛋白,产出错误的亲和力与功能数据,直接污染训练集。
Low-activity target protein in the wrong conformation produces wrong affinity and function data — and contaminates the training set directly.
多数服务止步于快速表达和 KD。没有细胞层面的功能验证,模型学不到真正的 MoA 与成药性规律。
Most providers stop at fast expression and KD. Without cell-level functional validation, a model never learns real MoA or developability rules.
单季度(90 天)可完成的验证轮次Validation rounds per 90-day quarter
按传统流程 8 天/轮的中值估算。对 AIDD 而言,这意味着回流进模型的实验数据量翻倍——不是快一点,是多一倍的训练信号。
Estimated at a midpoint of 8 days per conventional round. For AIDD this means twice the experimental data flowing back into the model — not marginally faster, but double the training signal.
02 流程Flow
自客户提交序列当天起算。省掉的是抗体纯化环节——传统流程里最耗时的一段。Counted from the day you submit your sequence. What we remove is antibody purification — the longest stretch of the conventional workflow.
03 闭环The loop
AIDD 不是一次性的筛选,是一个不断收紧的循环。每转完一轮,就有新一轮真实数据回流给模型;循环完成得越快,模型积累的真实数据就越多。AIDD isn't a one-shot screen; it's a loop that keeps tightening. Each completed turn sends a new round of real data back to the model — the faster each turn completes, the more real data the model accumulates.
生成模型一天产出成百上千条候选全抗序列,等待被真实世界打分。
A generative model produces hundreds of candidate full-IgG sequences a day, waiting to be scored by the real world.
合成、转染、分泌真实分子——从全长 IgG 到 VHH / 小蛋白,检测的都是最终形式本身,不是替身。
Synthesis, transfection, secretion of the real molecule — from full-length IgG to VHH or mini-protein, what is measured is the final format itself, not a proxy.
结构化记录直接进管线,成功与失败的样本都保留——负样本同样是训练信号。
Structured records go straight into the pipeline; hits and non-binders are both kept — a negative is training signal too.
上清直接上机,得到 KD、ka、kd 与拟合优度,附带原始曲线。
Supernatant goes straight onto the instrument, yielding KD, ka, kd and goodness of fit, with the raw trace attached.
快至四天转一圈。季度内多转 11 圈,模型多拿一倍的真实反馈——这是速度对 AIDD 唯一有意义的定义。
As fast as four days per turn. Eleven extra turns a quarter, and the model gets twice the real feedback — which is the only definition of speed that matters to AIDD.
04 交付物Deliverable
每一条设计对应一条结构化记录,含原始曲线、动力学参数与拟合优度。交付不止 KD——从序列到结合动力学全过程的数据(含原始数据)均可提供。字段字典随首次交付一并提供。Every design maps to one structured record carrying the raw trace, the kinetic parameters and the goodness of fit. Delivery goes beyond KD — data across the whole sequence-to-kinetics workflow, including raw data, is available. The field dictionary ships with the first delivery.
{ "design_id" : "Ab-002", "assay" : "BLI", "kd_nM" : 4.8, "ka_1/Ms" : 5.1e5, "kd_1/s" : 2.4e-3, "fit_r2" : 0.991, "rmax_nm" : 0.42, "raw_trace_csv" : "Ab-002_raw.csv"}
design_id,assay,kd_nM,ka_1/Ms,kd_1/s,fit_r2,rmax_nm,raw_trace_csvAb-001,BLI,13.2,2.4e5,3.2e-3,0.994,0.51,Ab-001_raw.csvAb-002,BLI, 4.8,5.1e5,2.4e-3,0.991,0.42,Ab-002_raw.csvAb-003,BLI, — , — , — , — ,0.03,Ab-003_raw.csvAb-004,BLI,27.6,1.8e5,4.9e-3,0.988,0.47,Ab-004_raw.csvAb-005,BLI, 0.9,7.6e5,6.8e-4,0.996,0.55,Ab-005_raw.csv
给的是测量事实与拟合指标,不是替你下的合格判定——阈值由你的管线来定。经验分析可作为额外数据交付:可按既定规则与算法赋值的高通量数据直接给出分析值,需人工判断的部分单独标注、置于标准数据包之外。
What ships are measurement facts and fit metrics, not a pass/fail verdict made on your behalf — the thresholds belong to your pipeline. Experience-based interpretation is available as an additional deliverable: rule- and algorithm-based analysis values where high-throughput data allows, and manually assessed items flagged separately, outside the standard data package.
| 字段 | 含义 |
|---|---|
| design_id | 与你提交时的命名一一对应 |
| assay | 检测方法 |
| kd_nM | 平衡解离常数 |
| ka_1/Ms | 结合速率常数 |
| kd_1/s | 解离速率常数 |
| fit_r2 | 拟合优度 |
| rmax_nm | 响应幅度,信号强弱的直接证据 |
| raw_trace_csv | 该样本的原始曲线文件(时间-响应两列) |
| Field | Meaning |
|---|---|
| design_id | Maps one-to-one to your own naming |
| assay | Measurement method |
| kd_nM | Equilibrium dissociation constant |
| ka_1/Ms | Association rate constant |
| kd_1/s | Dissociation rate constant |
| fit_r2 | Goodness of fit |
| rmax_nm | Response amplitude — direct evidence of signal strength |
| raw_trace_csv | Raw curve file for that sample (time / response columns) |
05 不止 KDBeyond KD
四天这一轮解决的是排序:谁值得往下做。进入 Leads 阶段后,纯化蛋白进入完整的功能与成药性验证——这是把「能不能结合」变成「能不能起效」的地方。
The four-day round solves ranking: which designs are worth taking further. At the Leads stage, purified protein enters full functional and developability validation — this is where does it bind becomes does it work.
理化 · 亲和力 · 初步成药性Biophysics · Affinity · Preliminary developability
MoA 验证 · 效应功能MoA validation · Effector function
hiPSC 来源 · 3D 高阶模型hiPSC-derived · 3D models
06 数据从哪来Where the data comes from
测得准的前提是靶点对。这些资源不是外采的,是同一家做出来的——所以数据前后一致、可追溯。Measuring accurately starts with the right target. These resources aren't bought in — they're built by the same company, which is why the data stays consistent and traceable.
逐批活性验证;HEK293 表达体系,与人体一致的糖基化与折叠;多标签、多物种、多形式覆盖。全球 Top 药企与 Biotech 广泛使用。
Lot-by-lot activity validation; HEK293 expression for human-like glycosylation and native folding; multi-tag, multi-species, multi-format coverage. Used widely by leading global pharma and biotech.
GPCR、离子通道、转运体、GPI 锚定与黏附分子。膜蛋白占药物靶点 40% 以上,却是公开数据最稀缺的一块。
GPCRs, ion channels, transporters, GPI-anchored and adhesion molecules. Membrane proteins are over 40% of drug targets — and the scarcest public data of all.
过表达株、KO 株与 300+ 报告基因功能株。同一株细胞可从早期筛选一路用到 CMC 质量放行。
Overexpressing lines, KO lines and 300+ reporter-gene functional lines. One line carries from early screening all the way through to CMC release.
07 边界Boundaries
不进入任何内部数据库,不用于我们自己的模型训练。
Sequences never enter any internal database and are never used to train our own models.
序列及由其产生的全部实验数据,所有权与知识产权归客户。
The sequence and all experimental data derived from it — ownership and IP remain with the customer.
覆盖中美欧主要市场,IND / BLA 申报无第三方 IP 隐患。
Covering the US, EU and China; no third-party IP exposure at IND / BLA filing.
百普赛斯集团(301080.SZ),13,000+ 客户。
ACROBiosystems Group (301080.SZ), 13,000+ customers.
08 开始Start
不用先谈合同。我们可以先发一份完整的样例数据包,或者用你手上已知 KD 的内部标准品做一轮小规模试运行——测出来的值和你手里的真值一对,准确性、报告质量、交付格式一次见分晓。
No contract needed first. We can send a complete sample data package, or run a small pilot on an in-house standard whose KD you already know — compare our number against your truth and accuracy, report quality and delivery format all settle in a single round.
CSV 数据包 + 字段字典 + 一条含原始曲线的完整记录,附一行命令转 JSON。看完再决定要不要谈。
CSV package + field dictionary + one complete record with its raw trace, plus the one-line JSON conversion. Decide after you've read it.
序列提交前完成 NDA,明确使用边界与数据所有权。
Executed before any sequence is submitted, defining use boundaries and data ownership.
建议夹带你已知 KD 的标准品做盲测,用你自己的标准验收我们。
We suggest spiking in a standard whose KD you know, blinded — accept us on your own criteria.