Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation
MarkTechPost
Read Full Article at MarkTechPost →Ad Slot — In-Article (728x90)
Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2. 24% to 1.
77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21. 2% […] The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost.
This is a summary. For the full story, read the original article at MarkTechPost.
Original source: MarkTechPost