aa37067f79
- Add reproducible v6 boundary, blessing, and consensus-adjudication corpora, plus the tiny-transformer trainer and v6 release-gate evaluator that gate every candidate on the deployed baselines. - Wire consensus-label merging, product-policy anchor evaluation, and sealed blessing benchmark review with their pytest coverage. - Refresh open-training corpus generation, iterative retraining runner, and random-holdout evaluation so v6 candidates can be benchmarked end-to-end.
1.3 KiB
1.3 KiB
Clipboard semantic evidence adjudication v1
Review only the fields listed in unresolvedFields. Judge from the text itself;
do not inspect previous model votes, source labels, or another adjudicator.
Return one JSON object per input record:
{
"id": "same id",
"resolutions": {
"task": "true",
"ambiguous": "false"
},
"confidence": {
"task": 0.97,
"ambiguous": 0.94
},
"evidence": {
"task": "send the report",
"ambiguous": "by Friday"
}
}
Requirements:
resolutions,confidence, andevidencemust contain exactly the fields inunresolvedFields.- Intent and flag values are
true,false, orunknown. - Sentiment values are
positive,neutral,negative, orunknown. - Evidence must be a short exact quote copied from the input text.
- Use
unknownwhen the text alone does not justify a decision. - Confidence is per field and must be between 0 and 1.
- Do not output explanations, markdown, or additional fields.
Use the boundaries from labeling-instructions-v2.md. In particular, distinguish
requests from personal plans, genuine information questions from request-shaped
commands, direct messages from terminal notices, and expressed wishes from
quoted, future, sarcastic, or received blessings.