
Prompt Adherence
A visually appealing generation can still be a poor response to the prompt.
Evaluators assess whether the output captures the requested subject, attributes, actions, environment, composition, spatial relationships, and creative direction. More complex prompts can also be reviewed for failures such as missing instructions, incorrect object counts, attribute confusion, or relationships being applied to the wrong subject.
This creates a clearer picture of whether a model is simply producing attractive content or actually following user intent.


























