m-hua-01/ai-psychosis ·
github.com · 568 words · saved by 1 readers
No description, website, or topics provided.
Automated Red Teaming for AI-Induced Psychosis Tim Hua and AIs A simple red teaming framework for testing how AI models respond to psychotic characters. Note: Grok-4 on openrouter has started refusing to role play as a psychotic user. However, I found that Grok-3 is just as good and doesn't refuse. Quick Start Installation Install the environment using uv : uv sync Basic Usage Run red teaming on all default models with all characters: uv run redteaming_systematic.py 📁 Project Structure Core Scripts redteaming_systematic.py - Main red teaming script with batch processing capabilities results_a
related reading
- Lakera – Test your AI hacking skillsgandalf.lakera.ai
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- Perplexityperplexity.ai
- Unrestricted AI API + Enterprise Policy Gateway | abliteration.aiabliteration.ai
- Hume AI - The AI toolkit for voice and emotionhume.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Prompting best practicesdocs.anthropic.com
- Datacurve | The data engine for frontier AIdatacurve.ai
- There's An AI For That® — The front page of AItheresanaiforthat.com
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts, Internal Tools & AI Modelsgithub.com
- Goodfire AIgoodfire.ai
- BalatroBenchbalatrobench.com