Is the problem in the source, instructions, or output?
Open the run from Runs and compare the exact source row with what the pinned job version produced.
- Open a representative unexpected row and read its source fields and complete output.
- Compare it with a correct row from the same run.
- Check whether required context was missing, ambiguous, or stored in a different column.
- Review the job instructions and output fields attached to that run.
How do I improve the result?
Return to the original thread in the workspace or start a new thread with the same dataset.
- Describe the failed expectation with concrete examples.
- Clarify the decision rule, required context, or output field.
- Ask Everyn to prepare a new sample proposal.
- Approve the new sample and compare it with the earlier run before considering a full list.
If a result contradicts the recorded source and instructions in a way you cannot explain, contact support with the run ID and affected row IDs.