4 min read

What would you actually do with that prediction?

By

Cover image generated with AI, guided by my detailed design instructions.
Cover image generated with AI, guided by my detailed design instructions.

So let’s say I have twenty enzyme candidates and enough money to test four. I run a model, get a ranking, and prepare to choose.

Taking the top four sounds reasonable. But suppose one of the lower-ranked candidates could help us understand something the model is uncertain about. Testing it might change how we interpret several other candidates.

Would that be worth one of our four experiments?

It depends. Are we trying to find something that works immediately? Understand why it works? Gather information that will make the next round of experiments more useful? We might care about all three, with different priorities.

A prediction helps us think through that decision. We still have to decide how to use it.

I build predictive models, I care about getting them right. That is why I keep coming back to this question. Once we say a model will improve scientific discovery, I want us to follow the claim a little further.

What does somebody actually do differently because the model exists?

Danaher’s announcement about its planned autonomous research lab caught my attention for this reason. The company describes a workflow connecting molecular design, experimental testing, and model updates, with operation at scale expected in early 2027. The projected speed improvements are still targets to be demonstrated.

The University of Maryland’s CRAB Lab project raises similar questions. It aims to let researchers program automated biomanufacturing experiments, with substantial work going into the data systems, experiment coordination, and models that help determine subsequent steps.

For me the interesting part is how these pieces will work together when the experiments start producing inconvenient results.

My background in control engineering makes that a familiar concern. You measure a system, make a decision, act, and observe the response. You also have to understand the measurement well enough to know what it is telling you.

Biology adds its own complications. If an enzyme appears inactive, was it actually inactive under those conditions? Was there a problem with the measurement? Did we record enough information to investigate the difference later?

Now imagine that result feeding automatically into the next round of model training.

A technical failure could become a lesson about biology that the model should never have learned. The software might run successfully through the entire process.

And honestly.. I think the people preventing that kind of mistake deserve more attention.

Keeping sample identities consistent, tracking experimental conditions, checking units, deciding how to represent failed experiments. Each requires judgment about what the data mean. Calling all of it “building the pipeline” can hide how much scientific understanding the work demands.

Automation makes these questions more pressing. If the measurement rewards an artifact, a faster loop could help us pursue that artifact more efficiently. We have to establish that the thing being optimized represents the thing we actually care about.

This also changes what I want to see in an evaluation.

If a model helps select experiments, how did those choices compare with a sensible alternative under the same budget? If it is supposed to save time, where was time actually saved? If the result surprised us, what did we change afterward?

And yes a carefully evaluated predictor can be a complete scientific contribution. So can a dataset, an experimental method, or a better way to represent a biological problem.

The evidence should reach as far as the claim. When we claim better decisions, we should show what happened to the decisions.

I think this question travels well beyond biology. A maintenance forecast, a recommendation for a software team, or a model guiding an engineering design eventually meets someone with limited resources who has to choose what to do.

That person needs to understand what the output supports, where it is uncertain, and when to ask for more evidence.

This is the part I want to keep developing in my own work. I enjoy building models. I’m also interested in the reasoning and engineering needed to make their results useful to the next person.

So when I read that a model will accelerate discovery, I find myself looking for one more figure.

Show me what happened when somebody acted on the prediction. I want to see what we learned from that, including the awkward results.