Move from scattered scientific evidence to SAR your team can use
Excelra finds the evidence that matters, turns complex patents and papers into traceable records, and delivers data aligned to your program and informatics environment. Start with mining, curation, or an existing GOSTAR slice to reduce manual search and cleanup while giving scientists a stronger basis for modeling and decisions.
Clear the data bottleneck holding up the next decision
Confirm that sufficient evidence exists, convert known sources into usable records, or license a GOSTAR slice that is already curated. Each route removes a different source of delay, and they can be combined when the program requires it.
Data mining
Avoid funding a low-yield dataset buildScientific search specialists work across patents, journals, and accessible grey literature, then assess each candidate against your inclusion criteria. You learn where extractable SAR exists, how dense it is, and whether the evidence can support the intended dataset.
Data curation
Reduce cleanup, rework, and hidden data errorsExcelra extracts and standardizes structures, targets, activity values, units, and assay context into your chemical entity model. You receive traceable records built for analysis, modeling, and integration, including complex modalities and patent layouts.
GOSTAR partial backfiles
Put existing curated coverage to workIf GOSTAR already covers your target, modality, disease area, chemotype, or other defined slice, Excelra can license that curated slice with perpetual non-commercial rights to use the delivered data, subject to agreed terms. Your team gets focused historical SAR without rebuilding the same evidence from source.
Mining vs curation vs GOSTAR partial backfile
Choose the route that answers the immediate need while preserving a path to the final dataset. Screen a corpus before curation, enrich an existing GOSTAR slice, or move customer-supplied sources directly into a defined schema.
| Dimension | Data mining | Data curation | GOSTAR partial backfile |
|---|---|---|---|
| Primary job | Determine whether extractable SAR exists, where it sits, and at what yield | Produce traceable, ingestion-ready rows mapped to your fields and ontologies | Acquire an already-curated GOSTAR slice when coverage and schema fit align |
| Best when | Coverage is unknown; feasibility, competitive intelligence, diligence, or whitespace analysis comes first | You need custom fields, novel ontologies, client sources, or modalities beyond a stock slice | Strong GOSTAR coverage for your slice, and the GOSTAR model is acceptable (or needs only light reshape) |
| Typical deliverable | Screened source inventory, inclusion disposition, evidence-yield assessment, and next-step recommendation | SDF, CSV, TSV, Excel, JSON, or custom schema with standardized entities, source provenance, and multi-tier QC | Defined SAR export (structures, targets, activities, assays) under agreed rights |
| Rights posture | Scoped engagement deliverable | Durable internal analytical asset under agreed terms | Perpetual non-commercial rights to use the licensed slice, subject to agreed terms |
| Pilot to start | Feasibility and source-availability assessment | Representative curation and schema-alignment benchmark | Coverage assessment and representative schema review |
| You provide | Scientific question; patents vs journals; decision the answer unlocks | Sample patent or field list / schema snippet | Slice definition (target, modality, disease area, chemotype, and so on); required fields; format |
The value is choosing the right work in the right order. Use existing GOSTAR coverage where it fits, establish evidence yield before commissioning a larger build, and reserve curation for the fields, relationships, and source material your workflow genuinely requires.
Bring the discipline behind GOSTAR to your data requirement
GOSTAR is Excelra’s structure–activity relationship database. It contains data on more than 10.6 million pharmacologically active molecules from over 5 million screened documents, including patent coverage from USPTO, EPO, and WIPO. Excelra developed the data model and curation process and continues to expand and operate the database.
Strengthen automation with scientific ground truth
Public datasets and internal extraction pipelines can accelerate the work, but difficult records still create hidden risk. Excelra supplies validated reference sets and production curation for patent tables, stereochemistry, R-groups, Markush context, and multi-assay layouts that require scientific judgment.
-
Scientific methods proven at database scale
Mining and curation draw on the same discipline behind GOSTAR: controlled vocabularies, standardized activity types and units, source-linked records, and multi-tier review of structures and assays.
-
Partial backfiles are GOSTAR slices
License historical SAR for an agreed slice with perpetual non-commercial rights to use the delivered data, subject to the applicable terms.
-
Scientists adjudicate scientific meaning
Domain experts apply inclusion criteria, resolve chemotype and assay context, and preserve the relationships among structures, targets, activities, and source evidence.
-
Modalities and deliverables your stack can use
Small molecules, TPDs/degraders, and large molecules when the work calls for them, delivered for on-prem or in-stack use under agreed terms.