Repeating a calculation, respecting a permission and lowering a cost are different achievements.

An accuracy percentage tells us little until we know what was counted. A system may have answered a question, calculated a viable alternative, detected missing information or prevented an action that was not authorised.
Presenting those outcomes as a single figure for “reliability” makes for a simpler presentation, but makes it harder to understand what to expect from the system. When I evaluate a proposed AI application in energy, I want to start with the problem it was meant to solve. The architecture and model come afterwards.
Suppose a system proposes lowering a facility’s costs by shifting some of its consumption to another time. One test might check that the multiplication of consumption by price is correct. A more complete test would ask whether that shift is feasible, whether it requires additional staff or whether it affects production.
Both tests are useful, but their scope is different. Before comparing results, I would define the objective, the available data and the conditions that must be respected. I would also make clear which tools each approach received and how much time and computing resources it could use.
Comparing a model on its own with another system that has documentation, an optimiser and business rules lets us study those configurations. To attribute the difference to a particular component, we would need to examine its contribution within the whole system.
I would also include conventional solutions that do not use language models. If a simple procedure solves the problem at lower cost and with sufficient quality, that is a relevant result. An evaluation should allow us to question the architecture we chose.
Determinism is valuable when we need to review a calculation. Getting the same result under the same conditions makes it easier to identify what has changed and examine the procedure. Even so, a formula can use the wrong unit and reproduce the same error perfectly.
For a forecast, we can preserve the data, model version and procedure used. That lets us review how the forecast was produced; its predictive ability is tested by comparing it with what happens. These are two different questions: can we reconstruct the result, and was that result appropriate for the problem?
I would include difficult cases: incomplete data, amended contracts, equipment with similar names or constraints that appear after an alternative has been calculated. I want to identify where the system stops being reliable and what it does at that point. Asking for an additional check may be better than completing the process using an unstated assumption.
In my research, I have addressed these questions from different angles. The energy architecture proposal describes one way to organise domain representation and decision-making. Energy Decision Benchmark sets out criteria for examining capabilities and operating conditions. SARA Scale proposes a classification focused on governance properties.
These are my own works, published as preprints, with different scopes. I present them as proposals that can be compared with other architectures and examined through independent evaluations.
A test can check that a revoked authorisation blocks an action. That tells us about system control. Another can examine whether the system finds a valid solution for a new combination of constraints. That tells us about its capability. We need both to understand it better.
Likewise, describing an architecture or an invention and demonstrating how an implementation behaves are different steps. To substantiate a specific capability, we need to identify what implements it, how it was tested and what results were observed. The same standard should apply to my own work.
For building control, BOPTEST provides emulators and metrics for comparing strategies under defined conditions. These environments let us examine proposals before transferring them to a facility, while keeping clear which cases the results refer to.
There is another step afterwards: checking what was approved, what was actually implemented and what happened. A recommended tariff may never be taken up. A lower bill may reflect both a contract change and reduced activity at the facility.
When communicating a result, I would distinguish between a calculated alternative, an action that was carried out and an observed effect. If we attribute an improvement to the decision, I would explain what it was compared against and which other conditions changed.
The information I would ask an evaluation to provide is specific: the cases used, versions, input data, success criteria, resources consumed and failures found. With that information, we can discuss the result, decide where to use the system and plan the next test.
Before asking how accurate an AI is, I want to know what problem it solves, what alternative it is measured against and what it does when information is missing.
Founder of TheryOS, business owner and entrepreneur in the energy and financial sectors.