Loading...
Loading...
Open-source GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL model from xw1234gan — available for download and self-hosting on Hugging Face.
Best for
Local run command
pip install transformers accelerate && python -c "from transformers import AutoModelForCausalLM, AutoTokenizer; tok = AutoTokenzier.from_pretrained('xw1234gan/GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL'); model = AutoModelForCausalLM.from_pretrained('xw1234gan/GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL'); print('Loaded xw1234gan/GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL')"Presumptive Specs:
Community reliability estimate · not official
About this score: Community-estimated based on user reports and publicly available benchmark data (e.g. TruthfulQA). This is not an official score from the model provider. Scores may be inaccurate — always verify with the official leaderboard before making production decisions.
Not enough historical data yet. Check back after the next pricing sync.
Provider
Learn how to use GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL
💡 Tip: Start with the sample prompts above to see how GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL works best.
Proven prompts shared by the community for this model