Move from scattered scientific evidence to SAR your team can use

Excelra finds the evidence that matters, turns complex patents and papers into traceable records, and delivers data aligned to your program and informatics environment. Start with mining, curation, or an existing GOSTAR slice to reduce manual search and cleanup while giving scientists a stronger basis for modeling and decisions.

Scientist-led screeningInclusion criteria guide every search
Three-tier scientific QCCuration, review, and senior QC
Patent-informed coverageUSPTO · EPO · WIPO
Schema-aligned deliverySDF · CSV · TSV · Excel · JSON · custom

Clear the data bottleneck holding up the next decision

Confirm that sufficient evidence exists, convert known sources into usable records, or license a GOSTAR slice that is already curated. Each route removes a different source of delay, and they can be combined when the program requires it.

Coverage is uncertain

Data mining

Avoid funding a low-yield dataset build

Scientific search specialists work across patents, journals, and accessible grey literature, then assess each candidate against your inclusion criteria. You learn where extractable SAR exists, how dense it is, and whether the evidence can support the intended dataset.

Sources are known; usable rows are not

Data curation

Reduce cleanup, rework, and hidden data errors

Excelra extracts and standardizes structures, targets, activity values, units, and assay context into your chemical entity model. You receive traceable records built for analysis, modeling, and integration, including complex modalities and patent layouts.

The required coverage may already exist

GOSTAR partial backfiles

Put existing curated coverage to work

If GOSTAR already covers your target, modality, disease area, chemotype, or other defined slice, Excelra can license that curated slice with perpetual non-commercial rights to use the delivered data, subject to agreed terms. Your team gets focused historical SAR without rebuilding the same evidence from source.

Mining vs curation vs GOSTAR partial backfile

Choose the route that answers the immediate need while preserving a path to the final dataset. Screen a corpus before curation, enrich an existing GOSTAR slice, or move customer-supplied sources directly into a defined schema.

Dimension Data mining Data curation GOSTAR partial backfile
Primary job Determine whether extractable SAR exists, where it sits, and at what yield Produce traceable, ingestion-ready rows mapped to your fields and ontologies Acquire an already-curated GOSTAR slice when coverage and schema fit align
Best when Coverage is unknown; feasibility, competitive intelligence, diligence, or whitespace analysis comes first You need custom fields, novel ontologies, client sources, or modalities beyond a stock slice Strong GOSTAR coverage for your slice, and the GOSTAR model is acceptable (or needs only light reshape)
Typical deliverable Screened source inventory, inclusion disposition, evidence-yield assessment, and next-step recommendation SDF, CSV, TSV, Excel, JSON, or custom schema with standardized entities, source provenance, and multi-tier QC Defined SAR export (structures, targets, activities, assays) under agreed rights
Rights posture Scoped engagement deliverable Durable internal analytical asset under agreed terms Perpetual non-commercial rights to use the licensed slice, subject to agreed terms
Pilot to start Feasibility and source-availability assessment Representative curation and schema-alignment benchmark Coverage assessment and representative schema review
You provide Scientific question; patents vs journals; decision the answer unlocks Sample patent or field list / schema snippet Slice definition (target, modality, disease area, chemotype, and so on); required fields; format

The value is choosing the right work in the right order. Use existing GOSTAR coverage where it fits, establish evidence yield before commissioning a larger build, and reserve curation for the fields, relationships, and source material your workflow genuinely requires.

Bring the discipline behind GOSTAR to your data requirement

GOSTAR is Excelra’s structure–activity relationship database. It contains data on more than 10.6 million pharmacologically active molecules from over 5 million screened documents, including patent coverage from USPTO, EPO, and WIPO. Excelra developed the data model and curation process and continues to expand and operate the database.

Strengthen automation with scientific ground truth

Public datasets and internal extraction pipelines can accelerate the work, but difficult records still create hidden risk. Excelra supplies validated reference sets and production curation for patent tables, stereochemistry, R-groups, Markush context, and multi-assay layouts that require scientific judgment.

  • Scientific methods proven at database scale

    Mining and curation draw on the same discipline behind GOSTAR: controlled vocabularies, standardized activity types and units, source-linked records, and multi-tier review of structures and assays.

  • Partial backfiles are GOSTAR slices

    License historical SAR for an agreed slice with perpetual non-commercial rights to use the delivered data, subject to the applicable terms.

  • Scientists adjudicate scientific meaning

    Domain experts apply inclusion criteria, resolve chemotype and assay context, and preserve the relationships among structures, targets, activities, and source evidence.

  • Modalities and deliverables your stack can use

    Small molecules, TPDs/degraders, and large molecules when the work calls for them, delivered for on-prem or in-stack use under agreed terms.