GitHub hosts HunyuanVideo, Tencent's open-source framework for large-scale video generation models, enabling AI-driven video creation.
IndiaSocialBench
About IndiaSocialBench
IndiaSocialBench assesses how language models handle emotionally complex Indian social scenarios such as indirect refusals, family negotiations, honor and shame dynamics, grief etiquette, and money between friends. The benchmark tests models in English, Hindi, and Hinglish using 18 pilot scenarios with 50 items per model across eight cultural dimensions. Each model’s score is accompanied by a transcript and judge explanation, allowing direct inspection of responses. The platform reports both overall performance and language-specific gaps, highlighting where models struggle when switching from English to Hindi. Scores range from 0 to 10, with lower scores indicating weaker cultural and emotional understanding. Reasoning-capable models are evaluated under capped and high-effort settings to ensure comparability, and refusals on family topics are tracked separately to avoid skewing results.
Key features
- Cultural dimension scoring (indirect speech, hierarchy, family, honor, rituals, money, support)
- Language mode switching (English, Hindi, Hinglish)
- Per-model transcript and judge explanation access
- Language gap analysis (English to Hindi performance delta)
- Refusal rate tracking on family-related topics
- Reasoning effort control (capped vs high)
- Bootstrap confidence intervals for scores
- Open benchmark with public dataset and methodology
Use cases
- Evaluating cultural and emotional intelligence of language models for Indian contexts
- Comparing model performance across languages in Indian social scenarios
- Identifying weaknesses in indirect speech and face-saving responses
Pros
- Tests culturally specific scenarios relevant to Indian social interactions
- Supports three language modes: English, Hindi, and Hinglish
- Provides per-model transcripts and judge explanations for transparency
- Compares performance across multiple cultural dimensions
- Tracks language gaps between English and Hindi responses
Cons
- Limited to 18 pilot scenarios and 50 items per model
- Reasoning effort is capped for most models, limiting high-compute comparisons
- No free access; requires a $30 API budget for evaluation
- Only evaluates models on Indian cultural contexts
Frequently asked questions about IndiaSocialBench
What does IndiaSocialBench measure?
IndiaSocialBench evaluates how language models handle emotionally complex Indian social scenarios such as indirect refusals, family negotiations, honor and shame dynamics, grief etiquette, and money between friends. It tests models in English, Hindi, and Hinglish across eight cultural dimensions.
Who should use IndiaSocialBench?
The benchmark is designed for developers, researchers, and organizations working with language models in Indian contexts. It helps assess whether models can navigate culturally nuanced conversations accurately.
How are model scores calculated?
Scores range from 0 to 10, with lower scores indicating weaker cultural and emotional understanding. Each model’s performance is accompanied by a transcript and judge explanation, and language-specific gaps are reported separately.
Does IndiaSocialBench support multiple languages?
Yes, the benchmark evaluates models in English, Hindi, and Hinglish, allowing comparison of performance across these languages. Language-specific gaps are highlighted to show where models struggle when switching languages.
How can I access IndiaSocialBench results?
Results are available on the IndiaSocialBench leaderboard, where users can view overall performance, language-specific gaps, and refusal rates. Each model’s score links to detailed transcripts and judge explanations.
What are the limitations of IndiaSocialBench?
The benchmark currently uses 18 pilot scenarios with 50 items per model, which may not cover all possible Indian social situations. Additionally, reasoning-capable models are evaluated under capped and high-effort settings to ensure comparability.