MIRA 2.0
Making a model's recommendation arguable
MIRA recommends operating changes in a molybdenum leach circuit. After onsite research, I designed a review flow that shows operators the forecast, current setting, and consequence of accepting, adjusting, or rejecting each recommendation.
- ⟦Ÿëáŕ ·⟧
- 2024—25
- ⟦Ḿÿ šçöþë · ·⟧
- Designed the operator-facing experience and information architecture; conducted onsite research at the plant. Data science and engineering owned the model and infrastructure.
- ⟦Çöľľáƀöŕáţöŕš · ·⟧
- Product team with data scientists, cloud engineers, plant metallurgists, and operations staff.
~20%
lower recommended fresh-ferric addition than the legacy bleed calculator
Monitored seven-day comparison, internal tracking.

A recommendation operators cannot challenge gets ignored
MIRA predicts impurity concentrations in the molybdenum leach circuit at Sierrita, outside Tucson, and recommends changes to throughput, fresh ferric addition, decant bleed, and digester runs. The open question was whether an operator on shift would act on it.
Operators already had a legacy bleed calculator and years of judgment. A screen that prints a number and offers a green button would ask them to replace both with blind trust. The design problem was to make the prediction arguable: show what the model expects under the recommendation, what it expects under the current setting, and where both sit against the spec limit.
The interface's job was not to be believed. It was to be verified.
Onsite, October 2024
I spent a day at the plant with operators and metallurgists. Two findings changed the structure of the product.
The first is that assays lag. Lab results arrive hours after the material they describe, so every recommendation rests on a slightly stale picture, and everyone on the floor already knows it. The second is that the decision is not made by one person at one moment. It crosses a shift handoff, and the person who acts is frequently not the person who reviewed.
Neither is visible from a model output or a stakeholder workshop. Both became structural. Data recency is printed on the home screen rather than implied. Any calculator run can be shared by URL or locked in as the committed strategy, with an author and timestamp attached, because the decision has to survive a shift handoff even when the reviewer changes.
The screen the product exists for
Review is a five-step flow with one recommendation type per step. The alternative was a dashboard asking for four decisions at once, which is how you get all four accepted without any of them being read.
On the fresh-ferric step, six impurity elements are drawn as small multiples on one shared axis under one shared spec line. Solid to the left of today, dashed to the right, which encodes the actual-to-forecast boundary with no legend to look up. The recommended setting and the current setting are both forecast, in two colors, on the same chart. When the current setting breaks spec and the recommendation does not, that reads as a shape before it reads as a number.
The interface offers three choices: accept the recommendation at its setpoint, adjust to a custom value, or reject and keep the current setting. Both setpoints are printed on the buttons, so the consequence of disagreeing is as legible as the consequence of agreeing. A metallurgist who wants to investigate rather than act can expand any chart or switch it to a table.
Trust as a first-class metric
The home screen carries two figures most ML products keep in a status deck: the share of recommendations operators approved, and the share of days anyone reviewed at all. Publishing a model's rejection rate to the people accountable for acting on it is uncomfortable and useful. A model nobody reviews is not a model in production; it is a model in a database.
The blend calculator's results screen follows the same instinct in reverse. It leads with the constraint violation—the number of days beyond target in red—above every positive number on the page. Then it shows the arithmetic underneath so the result can be checked rather than accepted.
What the evidence supports
Over one monitored week, internal tracking put MIRA-recommended fresh-ferric addition roughly 20% below what the legacy bleed calculator would have called for. The team's note said savings still needed month-by-month validation, so this site does not convert that comparison into an annual or dollar claim.
The home screen and blend calculator do not share the same chrome. The design system caught up mid-project, and I chose not to rework a surface already in use by people on shift. I would unify them now.
◇⟦Áŕţïƒáçţš · ·⟧

Step two of the recommendation wizard. Six impurities are forecast on a shared axis against one spec line, and accept, adjust and reject each show the setpoint it would produce. 
The feed blend input screen, pairing every candidate lot's selection slider with the assay chips needed to choose it, captured with a validation error, an active filter and a job toast all live. 
The job run result, leading with the schedule overrun before the production arithmetic, then eight range-bounded compliance meters that each state an in or out of spec verdict.A photographic avatar on the attribution chip is blurred.
◆⟦Đëçïšïöñš · ·⟧
⟦Ţĥë đëçïšïöñ · ·⟧
Make reject a full third option with its numeric consequence printed on it, rather than accept-or-dismiss.
⟦Ţĥë ţŕáđëöƒƒ · ·⟧
It puts disagreement one click away at the exact moment the program is being measured on adoption. A two-button flow would have shipped sooner and produced a better-looking approval rate.
⟦Ţĥë çöñšëqüëñçë · · ·⟧
Rejections became data instead of silence. Operators had somewhere to put disagreement other than not opening the app, and the team could see which recommendation types were being turned down and argue about why.
⟦Ţĥë đëçïšïöñ · ·⟧
Put six elements on one shared axis under one shared spec line, instead of six independently scaled charts.
⟦Ţĥë ţŕáđëöƒƒ · ·⟧
Elements with small absolute ranges get flattened and lose exactly the resolution a metallurgist wants. Covering that cost a per-chart expand and a chart-to-table toggle, both of which are extra surface to build, test and maintain.
⟦Ţĥë çöñšëqüëñçë · · ·⟧
Cross-element comparison is instant and a spec breach reads before any number is parsed. The people who needed precision got a route to it that is one interaction deep rather than a separate screen.
⟦Ţĥë đëçïšïöñ · ·⟧
Print the model's rejection rate and the staleness of the lab data on the operator home screen.
⟦Ţĥë ţŕáđëöƒƒ · ·⟧
Both facts argue against acting, on the screen whose entire purpose is to get someone to act. It also invited early questions from stakeholders about why the approval number was not higher, months before there was a good answer.
⟦Ţĥë çöñšëqüëñçë · · ·⟧
Review coverage became a tracked number rather than an assumption, and assay lag became a stated property of the system instead of an implicit limitation. The interface could not fix the lag, but it could make it visible.
Data visualizationMachine learningField researchInformation design