Maybe you should base this assessment on more than just vibes. Grok came out pretty balanced on independent assessments, where most other models were heavily biased: https://github.com/washingtonpost/political-bias-llm-eval/
Maybe you should base this assessment on more than just vibes. Grok came out pretty balanced on independent assessments, where most other models were heavily biased: https://github.com/washingtonpost/political-bias-llm-eval/