✳flâneur — a map of the web's best reading
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- GitHub - guidance-ai/guidance: A guidance language for controlling large language models. · GitHubgithub.com
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com
- Steering Black-Box LLMs with Advisor Modelsarxiv.org
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Meta-Prompt: A Simple Self-Improving Language Agentnoahgoodman.substack.com
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferencesarxiv.org
- GitHub - affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. · GitHubgithub.com
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com
- Chain-of-Thought Promptinglearnprompting.org
- llm-security/README.md at main · greshake/llm-security · GitHubgithub.com