{"id":155,"date":"2026-09-10T14:14:28","date_gmt":"2026-09-10T14:14:28","guid":{"rendered":"https:\/\/aiss.mohdluqman.com\/?p=155"},"modified":"2026-09-14T11:13:11","modified_gmt":"2026-09-14T11:13:11","slug":"gender-bias-in-text-to-video-generation-models","status":"publish","type":"post","link":"https:\/\/aiss.mohdluqman.com\/?p=155","title":{"rendered":"Gender Bias in Text-to-Video Generation Models"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>Mohammad Nadeem, Shahab Saquib Sohail, Erik Cambria, Bjorn W. Schluller, Amir Hussain<\/strong><\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><em>In IEEE Intelligent Systems, 2025<\/em><\/h3>\n\n\n\n<!--more-->\n\n\n\n<p class=\"wp-block-paragraph\">Humans build shared spatial understanding by communicating partial, viewpoint-dependent observations. We ask whether Multimodal Large Language Models (MLLMs) can do the same, aligning distinct egocentric views through dialogue to form a coherent, allocentric mental model of a shared environment. To study this systematically, we introduce COSMIC, a benchmark for Collaborative Spatial Communication. In this setting, two static MLLM agents observe a 3D indoor environment from different viewpoints and exchange natural-language messages to solve spatial queries. COSMIC contains 899 diverse scenes and 1250 question-answer pairs spanning five tasks. We find a capability hierarchy, MLLMs are most reliable at identifying shared anchor objects across views, perform worse on relational reasoning, and largely fail at building globally consistent maps, performing near chance, even for frontier models. Moreover, we find thinking capability yields gains in anchor grounding, but is insufficient for higher-level spatial communication. To contextualize model behavior, we collect 250 human-human dialogues. Humans achieve 95% aggregate accuracy, while the best model, Gemini-3-Pro-Thinking, reaches 72%, leaving substantial room for improvement. Moreover, human conversations grow more precise as partners align on a shared spatial understanding, whereas MLLMs keep exploring without converging, suggesting limited capacity to form and sustain a robust shared mental model throughout the dialogue<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<div class=\"wp-block-group is-layout-constrained wp-block-group-is-layout-constrained\">\n<div class=\"wp-block-group alignwide is-layout-constrained wp-block-group-is-layout-constrained\">\n<div class=\"wp-block-buttons is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-9b30a6b2 wp-block-buttons-is-layout-flex\">\n<div id=\"link\" class=\"wp-block-button is-style-outline is-style-outline--1\"><a class=\"wp-block-button__link has-light-green-cyan-color has-vivid-red-background-color has-text-color has-background has-link-color has-medium-font-size has-custom-font-size wp-element-button\" href=\"https:\/\/arxiv.org\/abs\/2603.27183\" rel=\"https:\/\/arxiv.org\/abs\/2603.27183\">PAPER<\/a><\/div>\n\n\n\n<div class=\"wp-block-button is-style-outline is-style-outline--2\"><a class=\"wp-block-button__link has-light-green-cyan-color has-vivid-red-background-color has-text-color has-background has-link-color has-medium-font-size has-custom-font-size wp-element-button\" href=\"https:\/\/scholar.googleusercontent.com\/scholar.bib?q=info:_q1HtpzcQHgJ:scholar.google.com\/&amp;output=citation&amp;scisdr=CoE6MIL1ELH51RbNLoo:AIVdB-wAAAAAaqfLNorLMw-n-bUrzl8_a5YoHu4&amp;scisig=AIVdB-wAAAAAaqfLNturYI-1NYsfgN31HEN8I7c&amp;scisf=4&amp;ct=citation&amp;cd=-1&amp;hl=en\"> CITE <\/a><\/div>\n<\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-group is-vertical is-layout-flex wp-container-core-group-is-layout-2c90304e wp-block-group-is-layout-flex\">\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Mohammad Nadeem, Shahab Saquib Sohail, Erik Cambria, Bjorn W. Schluller, Amir Hussain In IEEE Intelligent Systems, 2025<\/p>\n","protected":false},"author":8,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4],"tags":[],"class_list":["post-155","post","type-post","status-publish","format-standard","hentry","category-publications"],"_links":{"self":[{"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=\/wp\/v2\/posts\/155","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=\/wp\/v2\/users\/8"}],"replies":[{"embeddable":true,"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=155"}],"version-history":[{"count":37,"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=\/wp\/v2\/posts\/155\/revisions"}],"predecessor-version":[{"id":249,"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=\/wp\/v2\/posts\/155\/revisions\/249"}],"wp:attachment":[{"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=155"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=155"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/aiss.mohdluqman.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=155"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}