CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling
TL;DR - CellRFT is a reinforcement fine-tuning framework for single-cell perturbation models that directly optimizes non-differentiable biological evaluation criteria. It aims to align training with biologically meaningful outcomes rather than relying solely on surrogate losses.
- Uses policy-gradient optimization with evaluations of generated cell populations as training feedback.
- Combines multiple biological rewards through hierarchical reward aggregation.
- Experiments across different pretrained models report improved perturbation prediction.
- Results show that biological criteria can conflict or reinforce one another, highlighting implications for reward and evaluation design.