AI in orodja / Zapisek s terena
AI model izberite s preizkušanjem poteka dela, ki ga mora podpirati
Pred primerjavo konfiguracij modelov zgradite praktično ocenjevanje okoli svojih vhodov, meril pregleda in operativnih omejitev.
Ta zapisek s terena je objavljen v angleščini.
The most useful model comparison begins with the work, rather than a general ranking. A system that checks product records needs different evidence of quality from one that drafts support replies or reviews images. A compelling demonstration on an unrelated task does not settle that choice.
Treat the model as one part of a complete configuration: the instructions, input preparation, available tools, output format and review process. That is the system your users will experience, and it is the system worth testing.
Define success before comparing candidates
Write down what a correct result must contain and what would make it unacceptable. For a catalogue review, that might mean identifying the right record, citing the relevant field and leaving an unsupported correction unresolved.
For a support draft, the criteria might include answering the actual question, using the supplied policy accurately and avoiding promises that the available record cannot support. These are different tests, even if both workflows produce text.
OpenAI's evaluation guidance recommends defining an objective, collecting a relevant dataset, establishing metrics and comparing results as the system changes. It also warns against evaluation data that does not reflect the intended use. OpenAI evaluation best practices.
Use examples that reveal the difficult cases
Build a small, reviewed set of representative inputs. Include straightforward examples, incomplete information, contradictory records and cases where the correct outcome is to ask for more context.
Keep some examples aside while you refine the instructions. Otherwise, it is easy to tune the prompt around the same familiar cases and mistake that improvement for broader reliability.
Record the expected behaviour in plain language. You do not always need one exact reference answer, especially for writing tasks. You do need enough clarity for two reviewers to recognise the difference between a useful response and an unsupported one.
Compare actual configurations
Keep the input data and success criteria consistent, then record the model identifier, settings and prompt version used for each run. Product names alone are too broad to describe a reproducible test.
If one candidate receives different tools, a longer source document or a different review process, record that difference. The comparison may still be useful, but the result describes the whole configuration rather than an isolated model advantage.
Look at recurring failure types alongside overall quality. A system that usually writes polished responses but occasionally substitutes the wrong product identifier may be unsuitable for the task without an additional check.