ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
Abstract
ActReview is a rebuttal-guided post-training framework that generates diagnostic claims and concrete revision suggestions for peer review by leveraging author responses as latent supervision.
As LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identifies weaknesses but also guides authors toward concrete revisions. We study this as Actionable Peer-review Generation and decompose it into two subtasks: diagnostic claim generation and revision suggestion generation. We introduce ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans. Our central insight is that author rebuttals reveal plausible actions for addressing reviewer concerns and can therefore provide latent supervision for revision-oriented feedback. From real review-rebuttal threads on OpenReview, we construct ActReview-40K by aligning reviewer weaknesses with author responses and grounding the resulting feedback in localized paper evidence. We post-train Qwen3-8B-Base with multi-task supervised fine-tuning followed by GRPO using candidate-aware, weakness-specific rubric rewards. We also introduce ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness. Experiments show that ActReview outperforms prior specialized review-generation models on actionability and grounding while remaining competitive with strong prompt-based LLMs. Human evaluation confirms improved revision usefulness while revealing a remaining gap in technical accuracy, and additional analyses support generalization to held-out papers and robustness across independent judges.
Community
π Excited to share ActReview!
We ask: Can LLM-generated peer reviews not only identify weaknesses, but also help authors revise their papers?
Our key insight is that reviewer comments reveal what is wrong, while author rebuttals often reveal how the concern can be addressed. We use this signal to build ActReview-40K, aligning reviewer weaknesses with rebuttal responses and localized paper evidence.
We train ActReview with multi-task SFT followed by GRPO using weakness-specific rubric rewards, and introduce ActReview-Bench, a human-curated benchmark of 1,000 examples.
ActReview produces feedback that is more actionable and grounded than prior specialized review-generation models.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CriticGen: Generation-Aware Evaluation as Actionable Feedback (2026)
- Co-Evolving Actor-Conditioned Critics for Non-Verifiable Generation (2026)
- AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification (2026)
- RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing (2026)
- From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers (2026)
- Small Language Models as Judges for Rubric-Based Reinforcement Learning (2026)
- V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.09076 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper