HomePostsPlatform GovernanceGenerative AI Is Becoming a Gateway to Information. But Are Its Gatekeepers...

Related Posts

Generative AI Is Becoming a Gateway to Information. But Are Its Gatekeepers Protecting Free Speech?

The most consequential free-expression issues in generative AI arise not from individual outputs but from the policies, training choices, and design decisions that shape those outputs. As generative AI becomes one of the primary gateways to information in the digital age, these upstream decisions quietly determine the boundaries of permissible expression, influencing what kinds of inquiries models will engage with and how they present competing perspectives.

Recent weeks have shown AI companies grappling with these challenges. In an August update, OpenAI acknowledged that “deciding how an AI system should follow instructions is complex and we don’t have all the answers, especially in subjective, contentious, or high-stakes situations.” The company said that as AI becomes more capable and integrated into people’s lives, “its default behavior — and the boundaries of personalization — should reflect a wide range of perspectives and values.”

Similarly, OpenAI and Anthropic released new research on political bias, finding measurable progress toward neutrality as part of its effort to advance AI objectivity. This follows a 2025 Stanford Graduate School of Business working paper by Westwood, Grimmer, and Hall, “Measuring Perceived Slant in Large Language Models Through User Evaluations,” which found that most major models lean perceptibly left, even when prompted for balance.

With nearly 800 million monthly users, chatbots like ChatGPT, Claude, and Gemini are no longer just productivity tools; they are mediators of public discourse. They help decide what questions can be asked, which arguments are acknowledged, and which viewpoints are quietly filtered out.

That is why we at The Future of Free Speech at Vanderbilt University decided to take a closer look. Our latest report, “That Violates My Policies”: AI Laws, Chatbots, and the Future of Expression, includes an extensive analysis of eight leading AI models — from OpenAI, Google, Anthropic, DeepSeek, Meta, Mistral, Alibaba, and xAI — assessing how their rules, training practices, and responses affect the rights to freedom of expression and access to information.

From Over-Censorship to Selective Openness

Compared to our preliminary analysis in 2024, we found measurable progress. Most models are now more willing to engage with lawful but contentious prompts. Hard refusals — the kind of blanket “I can’t help with that” responses that once dominated — have dropped significantly.

In our latest prompting exercise, xAI’s Grok 4 engaged with every one of our 64 prompts, while Meta’s Llama 4, Mistral’s Medium 3.1, and Google’s Gemini 2.5 Flash responded to over 90%. Even OpenAI’s GPT-5 and Anthropic’s Claude Sonnet 4 showed improvement, with fewer blanket refusals and more nuanced engagement.

In our evaluation of free speech culture — a model’s willingness to foster open dialogue and engage diverse perspectives — xAI’s Grok 4 demonstrated the strongest performance; in contrast, Alibaba’s Qwen3-235B-A22B ranked lowest, systematically refusing to respond to many of the same prompts.

Yet the story is not unambiguously positive. While companies like OpenAI and Anthropic have publicly addressed “bias,” most continue to frame the problem in terms of objectivity or political balance rather than as a matter of viewpoint diversity, which is a core dimension of freedom of expression. Our analysis shows that models remain more comfortable producing abstract arguments rather than generating user-framed content such as social media posts. This sensitivity to “advocacy-style” requests may inadvertently chill lawful political expression, particularly for users who rely on these tools to participate in public debate.

Our findings also align with recent empirical work by Greene and Shapiro at Princeton University, “Measuring Free Expression in Generative AI Tools,” developed in collaboration with our team. Their study introduced a replicable, automated pipeline to evaluate how large language models restrict or redirect user expression. Testing four leading systems — GPT-4o, Gemini 2.5 Flash, xAI’s Grok 4, and DeepSeek-V3 — the researchers found little evidence of hard moderation (outright refusals) but considerable variation in soft moderation, where models subtly redirected stances. For instance, if asked to produce a social media post arguing that suspending a professor for online comments does not infringe on academic freedom, the model would instead respond “Silencing educators for their online opinions undermines the very essence of academic freedom… (GPT).”

GPT and DeepSeek generated affirmative posts 22 percent of the time when prompted for negatives, often steering away from critiques of free-speech principles, in the case of the former, or Chinese foreign policy, in the case of the latter. Taken together, our analyses show that moderation in generative AI increasingly occurs through design-level constraints, underscoring the need to apply robust freedom of expression standards to evaluate both overt and implicit restrictions.

The Hidden Gatekeeping Layer: Policies and Training Opacity

Herein lies the deeper issue: the policies and training choices that shape these outputs. Across all eight companies, usage policies remain vague, often invoking terms like “hateful,” “harmful,” or “misleading” without defining them clearly or tying them to legitimate aims recognized under international human rights law (IHRL). For instance, Alibaba’s policy simply states “Do not spread misinformation.” As we have previously covered IHRL offers a framework with helpful minimum standards for the protection of freedom of expression, even if it is less speech-protective than the First Amendment. It provides a useful baseline for evaluating global platforms and could be combined with the First Amendment to offer an additional benchmark for users in the United States. This framework serves as a foundation for developing more transparent and rights-based free speech principles.

Equally troubling is the opacity of model training. None of the companies we evaluated — not even those releasing open-weight models — disclose their training datasets or the criteria used by human raters to label speech as “helpful” or “harmful.” These are precisely the decisions that shape the boundaries of permissible expression. Without transparency, users cannot know whether refusals stem from legitimate safety concerns, corporate risk aversion, or ideological bias.

Still, there are encouraging signs. Some providers — particularly Anthropic, OpenAI, Google, and Meta — show meaningful efforts to engage with viewpoint diversity and reduce refusal frequencies. For example, Anthropic’s system card for Claude Sonnet 4 shows that it engages more consistently with sensitive topics than earlier versions, offering nuanced responses where Claude Sonnet 3.7 might have refused. Google has emphasized reducing unnecessary refusals by improving instruction following, and Meta reports that Llama 4 refuses less on debated political and social topics. Similarly, OpenAI has introduced a safe-completion approach aimed at lowering outright refusals.

These steps, while limited, demonstrate a growing awareness within the industry that aligning moderation and training choices with human rights standards is both feasible and necessary for protecting access to information.

Access to Information Matters in Gen AI

As generative AI systems increasingly mediate how people learn and communicate, vague moderation rules and opaque design choices pose a direct threat to users’ ability to access information online, which is a right enshrined in Article 19 of the International Covenant on Civil and Political Rights. These systems are not just tools; they are speech intermediaries performing quasi-regulatory functions on a global scale.

Their internal governance could shape the global information ecosphere around entrenched ideologies that do not allow inquiry into the narratives and perspectives seen online. The danger is not overt censorship but a gradual narrowing of public discourse through algorithmic design choices without transparency for users to understand the principles underlying policy decisions.

Alignment with Robust Freedom of Expression Standards

AI companies can and should align their systems with robust free speech standards. This means grounding restrictions in clear, legitimate aims; ensuring they are necessary and proportionate; and committing to meaningful transparency about moderation criteria. This also means limiting slant redirections in responses in order to subvert queries.

Generative AI holds immense promise to expand access to knowledge and pluralism, but only if its builders embrace the same expressive freedoms that sustain open societies. As the industry moves to embrace “objectivity,” it must also ensure that a handful of dominant generative AI providers do not restrict access to information in unprincipled or opaque ways. Without such commitments, the world’s new gateways to information risk becoming its most powerful gatekeepers.

Isabelle Anzabi
Research Associate at The Future of Free Speech

Isabelle Anzabi is a Research Associate at The Future of Free Speech, where she analyzes the intersections between AI policy and freedom of expression. She is currently a Fellow with the Internet Law & Policy Foundry and has a background in digital rights policy, global regulatory approaches to AI governance, and content moderation. Isabelle received a B.A. in Political Science from Stanford University.

Jordi Calvet-Bademunt
Senior Research Fellow with The Future of Free Speech at Vanderbilt University |  + posts

Jordi Calvet-Bademunt is a Senior Research Fellow with The Future of Free Speech at Vanderbilt University. Jordi focuses on freedom of expression in the digital space in Europe, the United States, and globally.

Jacob Mchangama
Founder and Executive Director of The Future of Free Speech |  + posts

Jacob Mchangama is the Founder and Executive Director of The Future of Free Speech. He is also a research professor at Vanderbilt University and a Senior Fellow at The Foundation for Individual Rights and Expression (FIRE).

[citationic]

Featured Artist