• Dataset Overlap Analysis
  • KKB per-kinase activity classifiers
  • KKB per-kinase potency regressors
  • KKB model performance rises with depth of curated data
  • Kinase Knowledgebase Interface
  • KKB Sample Data
  • PDF Processing Workflow
  • Data Validation Process
  • Curation Search Interface
  • 3D Protein Structure
  • Protein-Ligand Binding
  • Oncology Knowledgebase
  • OKB Sample Data
  • Protein Model
  • ML Training Datasets

Advancing Drug Discovery Through AI-Powered Solutions

Eidogen-Sertanty is dedicated to improving healthspan, medicine, and general well-being through cutting-edge pharmaceutical research tools.


LEARN MORE

Our Products

Kinase Knowledgebase: growth to more than 3 million biological activity data points Four-panel performance summary of the KKB per-kinase activity classifiers: ROC curves grouped by training-set size, the distribution of ROC-AUC across 392 kinases, accuracy rising with the number of measured compounds, and mean AUC by data-size band.

Per-kinase model performance across 392 kinases

Kinase Knowledgebase (KKB)

Currently the Kinase Knowledgebase Q2 2026 Release includes the following data:

  • Journal articles and patents: 10,326
  • Number of Biological Activity Data Points: 3,336,903
  • Number of unique kinase molecules with annotated assay data: 509,691
  • Number of all unique kinase molecules from patents and articles (with or without bio-activity data): 854,437
  • Number of unique kinase targets with assay data: 579
  • Number of annotated assay protocols: 105,916
  • Machine-learning models built from this release’s SAR: activity classifiers across 392 kinase targets, median ROC-AUC 0.91 (0.95 for well-studied kinases) – view model performance report

To show what this depth of curation supports, we build machine-learning models from the KKB SAR and report how they perform: an activity classifier for each of 392 kinase targets and a potency regressor for 319 of them. The panel at left summarises classifier accuracy — ROC curves grouped by how much data each kinase has, the spread of ROC-AUC, and how accuracy rises with the depth of measured chemistry.

Accuracy is highest for compounds chemically related to what a target already has in KKB and declines for novel scaffolds, so every prediction is reported with a similarity score against the model’s own training set. Performance is also measured against published data absent from the knowledgebase.

Search the Kinase Knowledgebase →
OKB

Oncology Knowledgebase (OKB)

Currently the Oncology Knowledgebase Q2 2023 Release includes the following data:

  • Journal articles and patents: 977
  • Number of Biological Activity Data Points: 146,206
  • Number of molecules from patents and articles: 66,198
  • Number of unique oncology targets with assay data: 1,158
  • Number of annotated assay protocols: 5,614
  • Number of disease models: 137
Learn More →
Two-panel diagram. Panel one: a compound, with its assay, primary target and potency, is reduced by a one-way SHA-256 hash to an irreversible encoded fingerprint, and the original structure is kept private. Panel two: two parties each hold a private list of encoded fingerprints, compare them, and learn only the percentage that overlaps.

Illustrative example. Figures shown are for explanation only and are not results from our datasets.

Dataset Overlap AnalysisNew

How much of a dataset do you already have? Our toolkit answers that for any two SAR datasets — any target family — without either side revealing a structure. Each party encodes its own data locally into irreversible SHA-256 fingerprints; only fingerprints are compared, and only the overlap figure is shared.

How it works → Download the toolkit →
TIP

Target Informatics Platform (TIP™)

The Target Informatics Platform's content grows weekly with more than 200K high resolution protein structures, ~1.2M annotated co-complex sites, and an additional ~875K predicted sites. TIP contains more than 710K protein chains representing every major drug target family.

Learn More →