AI for R&D Data: Connect Spreadsheets, PDFs, Lab Notebooks, and Experiments

Jonathan Woo
Jonathan Woo Chief Product Officer, ChemCopilot LinkedIn →

Chief Product Officer at ChemCopilot. Former VP of Product at Noble.AI (Science-Based Chemical AI), Co-founder/CTO at Nanostellar (Quantum Simulation & Catalysts), and NASA/Harvard ACIS Software Team Leader. Over 25 years of engineering experience pioneering AI-driven inverse formulation design, quantum materials modeling, and enterprise SaaS.

Chemical R&D teams rarely have all their knowledge in one place.

Experimental data may be distributed across Excel spreadsheets, PDFs, formulation records, laboratory notebooks, reports, and different databases. Some information is structured, while other valuable observations exist only as handwritten notes or documents.

ChemCopilot helps connect this fragmented R&D information so scientists, chemical engineers, and R&D Managers can search, analyze, model, and use their experimental history in one AI-powered environment.

Using AI, LLMs, data analytics, machine learning, and specialized chemical agents, ChemCopilot can turn historical R&D data into a foundation for better experiments and better candidates.

Can AI connect spreadsheets, PDFs, and laboratory notebooks?

Yes. ChemCopilot is designed to bring together fragmented R&D information from spreadsheets, PDFs, experimental records, and laboratory notebooks, making it easier for R&D teams to search, analyze, and use their historical data.

The objective is not simply to store files in one location.

The objective is to make the information inside those files usable for analysis, modeling, and decision-making.

For a chemical R&D organization, this can mean connecting information such as:

  • Formulation compositions

  • Ingredient quantities and ratios

  • Experimental results

  • Process conditions

  • Material properties

  • Test results

  • Historical formulations

  • Technical documents

  • Laboratory observations

  • Molecular structures

  • SMILES representations

  • Manufacturing parameters

Once these sources are connected, researchers can begin asking questions across their experimental history rather than searching through individual files.

The problem with fragmented chemical R&D data

Many R&D organizations have years of valuable experimental knowledge but still struggle to use it efficiently.

One researcher may have a spreadsheet containing formulation experiments.

Another may have a PDF containing test results.

A laboratory notebook may contain observations about why an experiment succeeded or failed.

A separate database may contain process conditions.

And the relationship between all of these sources may exist only in the knowledge of individual researchers.

This creates a common problem:

The organization has data, but the data is not necessarily connected.

As a result, teams may repeat experiments, spend time manually cleaning historical data, or make decisions without fully leveraging previous work.

AI can help change this by creating a layer between fragmented information and the researchers who need to use it.

Using AI and LLMs to create a connected R&D data layer

ChemCopilot uses AI and LLM technologies to help transform fragmented information into a more structured analytical environment.

One important part of this approach is creating an OLAP-style analytical layer over R&D data.

Instead of treating every spreadsheet, document, or experiment as an isolated object, the information can be organized around the relationships that matter to chemical R&D.

For example:

Ingredients → Ratios → Process Conditions → Experiments → Properties → Results

This allows teams to analyze their data across multiple dimensions.

A formulation scientist might want to investigate how changing an ingredient ratio affected performance.

A process engineer might want to understand how temperature or curing time influenced the result.

An R&D Manager might want to identify which experiments produced the strongest candidates and where the organization has already invested significant experimental effort.

The value comes from connecting these questions to the underlying data.

What about information in laboratory notebooks?

Not all R&D knowledge exists in spreadsheets.

Laboratory notebooks can contain some of the most valuable information in an organization, including observations that were never entered into a structured database.

Researchers may write comments about:

  • Unexpected reactions

  • Appearance or viscosity

  • Processing difficulties

  • Material behavior

  • Experimental failures

  • Changes made during an experiment

  • Observations that explain why a formulation worked or failed

Historically, this information has been difficult to use computationally.

ChemCopilot is preparing an OCR capability for laboratory annotations, designed to read information from scanned or handwritten laboratory notes and make that information available for analysis.

This capability is ready to launch.

The long-term opportunity is significant: instead of leaving years of experimental knowledge trapped in notebooks, organizations can begin incorporating those observations into their broader R&D knowledge base.

Analyze your R&D data with specialized AI agents

Once the data is connected, the next question is:

What can you actually do with it?

ChemCopilot uses specialized AI agents to help R&D teams analyze their experimental information.

Researchers can investigate relationships between variables such as:

  • Ingredients

  • Concentrations

  • Ratios

  • Materials

  • Process parameters

  • Experimental conditions

  • Molecular characteristics

  • Performance targets

AI can help researchers explore the data and identify relationships worth investigating.

At the same time, visualization tools can make complex experimental datasets easier to understand.

Instead of looking at thousands of rows in a spreadsheet, researchers can explore patterns visually and interact with the underlying information.

This combination of AI agents + connected data + visualization helps turn R&D data into something researchers can actively work with.

Can ChemCopilot analyze historical experiments?

Yes. Historical experiments can become the foundation for machine learning models that help R&D teams understand relationships between experimental inputs and outcomes.

This is particularly valuable for organizations with years of accumulated experimental data.

Imagine a formulation team with thousands of historical experiments.

Each experiment may contain information about:

Formulation + Process + Experimental Conditions → Performance

Instead of treating those experiments only as historical records, ChemCopilot can help R&D teams use them to build machine learning models.

The model can learn relationships between the variables used in previous experiments and the measured outcomes.

This creates a new possibility:

Your historical experiments can help determine what to test next.

Build machine learning models from your R&D data

ChemCopilot allows teams to create machine learning models using experimental datasets.

For example, a team developing a new formulation could model a target property such as:

  • Compressive strength

  • Viscosity

  • Stability

  • Cure performance

  • Thermal properties

  • Mechanical performance

  • Chemical performance

The model can incorporate formulation and process variables to understand how they influence the target.

This makes it possible to move from:

“What happened in our previous experiments?”

to:

“What does our experimental history suggest we should try next?”

That distinction is important.

The goal of machine learning in R&D is not simply to predict a number. It is to help researchers make better decisions about where to spend their next experimental effort.

Sweep ingredients, ratios, and process conditions

Once a predictive model has been created, R&D teams can explore the design space more systematically.

For example, a formulation model can investigate combinations of:

  • Ingredient concentrations

  • Component ratios

  • Additive levels

  • Process temperatures

  • Mixing conditions

  • Curing conditions

  • Reaction parameters

Instead of manually evaluating every possible combination, teams can use modeling to identify promising regions of the experimental space.

This can help researchers prioritize candidates before committing laboratory resources.

The workflow becomes:

Historical experiments → ML model → Design space exploration → Candidate selection → New experiments

This does not eliminate laboratory experimentation.

It makes experimentation more targeted.

Can molecular structures be included in R&D models?

Yes. Molecular structures can also become part of the modeling workflow.

ChemCopilot can work with molecular representations such as SMILES, allowing chemical structure information to be incorporated into analytical and machine learning workflows.

This creates an opportunity to connect molecular information with formulation and experimental data.

For example:

Molecular structure + formulation + process conditions → predicted performance

A team could investigate whether changes in molecular structure correlate with a target property, or use molecular information to evaluate potential candidates.

This becomes particularly interesting when molecular design and formulation development are connected rather than treated as completely separate workflows.

From molecular candidates to better formulations

Consider a formulation development workflow.

A team has historical experiments containing:

  • Existing ingredients

  • Ratios

  • Process conditions

  • Performance results

The team also has candidate molecules represented using SMILES.

The molecular information can be incorporated into the modeling workflow alongside formulation and experimental variables.

The resulting models can then help evaluate potential candidates against the desired target.

This creates a workflow that can move from:

Molecule → Formulation → Process → Prediction → Candidate

Rather than simply asking which molecules look interesting chemically, researchers can evaluate them in the context of the actual performance objectives of the formulation.

That can help create better candidates and reduce the amount of trial-and-error involved in development.

An example of an AI-powered R&D workflow

A connected ChemCopilot workflow could look like this:

1. Connect your existing R&D information

Bring together spreadsheets, PDFs, experimental records, formulation data, and other relevant sources.

2. Structure the information

AI and LLM technologies help organize fragmented information into a connected analytical layer.

3. Search and explore your experiments

Researchers can investigate historical experiments and identify relationships across formulations, materials, processes, and results.

4. Analyze with AI agents

Specialized agents help investigate the data and answer questions about the experimental history.

5. Visualize the results

Interactive visualization tools help researchers understand patterns and relationships in the data.

6. Build machine learning models

Use historical experiments to model relationships between inputs and target properties.

7. Explore the design space

Sweep ingredients, ratios, process conditions, or other variables to identify promising combinations.

8. Incorporate molecular candidates

Use molecular structures and SMILES as additional inputs to the modeling workflow.

9. Identify better candidates

Use the models to prioritize formulations or molecules that are more likely to meet the desired target.

10. Run targeted experiments

Take the most promising candidates back to the laboratory and generate new experimental data.

The new results can then become part of the organization's growing R&D knowledge base.

The goal is not to replace the scientist

AI-powered R&D should not mean removing scientists from the development process.

The scientist remains responsible for understanding the chemistry, defining meaningful objectives, interpreting results, and deciding which experiments are scientifically relevant.

AI can help with another part of the problem:

making the organization's accumulated experimental knowledge easier to access and use.

Instead of spending hours searching through spreadsheets and documents, researchers can spend more time investigating hypotheses.

Instead of testing every possible formulation combination, they can prioritize promising candidates.

Instead of treating historical experiments as archived information, they can use them as training data for new models.

Why connected R&D data matters

The real value of an R&D data platform is not simply having all your files in one place.

It is creating relationships between the information inside those files.

A spreadsheet may tell you what formulation was tested.

A laboratory note may explain what happened during the experiment.

A PDF may contain the resulting performance data.

A molecular structure may explain something about the chemistry.

A machine learning model can then help connect these variables to the target you are trying to optimize.

When these elements work together, historical R&D data becomes more than documentation.

It becomes an input to innovation.

Frequently Asked Questions

Can ChemCopilot connect laboratory spreadsheets and PDFs?

Yes. ChemCopilot is designed to connect different sources of R&D information, including spreadsheets, experimental datasets, PDFs, and other research records, so teams can work with their information in a more unified environment.

Can AI search my historical R&D experiments?

Yes. ChemCopilot uses AI and LLM technologies to help researchers explore and analyze connected R&D information, making historical experimental knowledge easier to find and use.

Can ChemCopilot read laboratory notebooks?

ChemCopilot is preparing an OCR capability for laboratory annotations that can extract information from scanned or handwritten laboratory notes. This capability is ready to launch.

Can I build machine learning models from my old experiments?

Yes. Historical experimental datasets can be used to create machine learning models that learn relationships between formulation, process, molecular, and performance variables.

Can ChemCopilot optimize chemical formulations?

Yes. Machine learning models can be used to explore ingredients, ratios, process conditions, and other variables to identify promising formulation candidates for a defined target.

Can molecular structures be used in the models?

Yes. Molecular representations such as SMILES can be incorporated into modeling workflows, allowing molecular information to be analyzed alongside formulation and experimental variables.

Does AI replace laboratory experiments?

No. The objective is to make experiments more targeted. AI and machine learning can help prioritize promising candidates, while laboratory experiments remain essential for validation and scientific discovery.

Turn your R&D history into a tool for your next experiment

Chemical companies have accumulated enormous amounts of experimental knowledge.

The challenge is making that knowledge accessible.

Spreadsheets, PDFs, notebooks, formulations, molecular structures, process data, and experimental results can contain valuable information about what has already been tried and what might work next.

ChemCopilot brings these elements together with AI, LLMs, specialized chemical agents, data analytics, machine learning, molecular representations, and visualization tools.

The result is a more connected approach to chemical R&D:

Connect your data. Understand your experiments. Build models. Explore possibilities. Design better candidates.

Because your previous experiments should not only tell you what happened.

They should help you understand what to try next.

Next
Next

The Best AI for Chemistry in 2026 -Top Tools Transforming the Field