Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
slime is an LLM post-training framework for RL Scaling.
Python 8.4k 1.2k
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Python 10k 1k
[ACL25' Findings] SWE-Dev is an SWE agent with a scalable test case construction pipeline.
Python 66
[ICLR'25] Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
Python 4
There was an error while loading. Please reload this page.