Tools
Our reward design combines correctness, preference, and efficiency. Preference only counts when the answer is correct. This keeps the model …
Perplexity's reward design system prioritizes correctness as a foundational requirement, then evaluates user preference and efficiency only when answers meet correctness standards. This hierarchical a
Perplexity's reward design system prioritizes correctness as a foundational requirement, then evaluates user preference and efficiency only when answers meet correctness standards. This hierarchical approach ensures the model optimizes for accuracy first before considering secondary factors like user satisfaction and computational efficiency.
Related
- This pipeline is why the same base model produces more accurate, better-cited, and more efficient answers inside Perplexity than out of the …
- Our researchers are heading to ICLR with new work: model efficiency, long-context reasoning, next-gen attention and decoding, and more. Chec…
- I'm releasing the 34 slides on how we design and train best-in-class edge models at @liquidai I presented these slides yesterday at @aiDotEn…
Source: Perplexity (X) | 2026-04-22