RL Post-Training on Macs
news
pluralis.ai
IrregularChat: AI & Autonomy
Jul 16
The post describes a multi-turn reinforcement learning run of LFM2.5-8B-A1B, an 8.3B-parameter mixture-of-experts model, using 14 consumer Macs distributed across four countries for rollout generation