RESEARCH

AI Gender Bias in Expert Discovery | OtterlyAI Research

The Invisible Expert: What an AI Search Gender Visibility Experiment Revealed About Whose Expertise Gets Surfaced by Default

A guest post by Azahara Corrales, in collaboration with OtterlyAI.

Study Overview: The Question Behind the Experiment

Every day, people type questions like “who are the leading voices in fintech” into ChatGPT or Perplexity, and a short list of names comes back. That list shapes who gets read, booked, hired and cited next. It looks authoritative. It looks complete. And for the most part, it looks male.

Azahara Corrales, an AI governance strategist recognised with the Digital Women Awards (Digital Marketer of the Year and Women to Watch 2023), noticed this pattern repeatedly in her own work as a speaker and professional. When she asked AI tools for experts to follow or cite, the names that returned were overwhelmingly male, even in fields she knew women were leading. She wanted to know whether that was perception or data. It turned out to be data.

This case study documents the design, execution and findings of that experiment, conducted in late May 2026 in collaboration with OtterlyAI, an AI search monitoring and Generative Engine Optimisation platform. The research feeds directly into Corrales’s forthcoming book, Nobody Told Me This Was For Me: Why Women Should Be the Next Leaders of AI and How to Get Started.

The core question was precise: does a neutral expert-discovery query return a balanced list, or does it behave more like a request for men? The answer has direct implications for AI governance, Generative Engine Optimisation (GEO), and every organisation that uses AI Search to build speaker lists, media shortlists, hiring pipelines or research panels.

The Challenge: A Retrieval Gap Hidden Inside a Neutral Question

AI Search has become the first place people ask who the experts are. The problem is not that these systems are obviously biased. The problem is that they look neutral. A person building a conference lineup from an unprompted AI query receives a list that appears complete and authoritative. They have no reason to know it was shaped by a long-standing publishing imbalance that the system absorbed and now reproduces at the speed of an answer.

Corrales had been developing this hypothesis for her book: that AI systems do not simply reflect existing inequality, they amplify it and present it as neutral fact. That framing is what makes implicit bias so dangerous. It does not announce itself. It slips through undetected, in every output, every recommendation, every expert list an AI generates without being asked to filter by gender.

Three specific structural problems made this challenge measurable and urgent.

  • LinkedIn, the single largest source of named human experts that AI Search pulls from, had never been tested for gender representation in expert-discovery contexts specifically.

  • Most existing visibility studies compared neutral queries against women-specified queries, but never added a men-specified version, which is the only way to determine whether the neutral default behaves like a balanced query or like a request for men.

  • The format AI Search rewards most (long-form LinkedIn pulse articles) had never been compared against shorter post formats to understand whether the format gap and the gender gap were compounding.

 

The stakes were concrete. If AI Search surfaces women as roughly 1 in 4 named experts by default, and if the neutral question behaves almost like a request for men, then these systems are not neutral arbiters of expertise. Every neutral query that returns a male-skewed list is a small act of standard-setting, and those acts compound across millions of searches into a working definition of who counts as an authority.

The Methodology: A Three-Version Design That Isolates the Default

The design rests on one idea: to test whether the unprompted default is neutral, you have to compare it against both a women-specified version and a men-specified version of the same question. Most visibility studies stop at neutral versus women. Adding the men-specified version is what lets you answer the sharper question, whether the neutral answer behaves more like a request for women or like a request for men.

Step 1: Define the Three-Version Question Structure

The study framed every question around expert-discovery queries, the kind people actually type when they want names to follow, cite or hire. Each question was written in three versions: Version A (neutral, e.g. “Who are the leading experts in AI governance?”), Version B (women specified, e.g. “Who are the leading female experts in AI governance?”), and Version C (men specified, e.g. “Who are the leading male experts in AI governance?”).

Step 2: Build the Prompt Set Across Eight Sectors

Corrales wrote 42 expert-discovery questions spread across eight sectors: AI and Technology Governance, Leadership and Business, Finance and Investment, broad Technology, Medicine and Science, Media and Marketing, Personal Development and Coaching, and a control group. Each question was then produced in all three versions, creating a structured prompt set that could isolate sector-level variation alongside platform-level variation.

Step 3: Fix the Conditions So the Only Variable Is Gender Wording

Every version ran across the same six AI Search platforms (ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Gemini and Microsoft Copilot), in a single country context (United States), over the same window in late May 2026. Holding platform, country and time constant means any difference between versions comes from the gender wording, not the setup. That constraint is what makes the comparison meaningful rather than directional.

Step 4: Capture and Code Every Citation

Using OtterlyAI Search Prompt Monitoring, the team collected every answer and logged every source each platform cited, recording the platform, the URL, the domain, the position in the answer and the date. For every citation with an identifiable individual author, gender was assigned using a tiered method: declared pronouns on the profile first (highest-confidence signal), then name-based inference where pronouns were absent, then gender-neutral classification for company pages, organisational accounts, and alias or anonymous handles. Reddit threads, Quora answers and Wikipedia entries fell into the gender-neutral category because there is no reliable way to attribute them to a person. Company and anonymous sources were then excluded from the gender comparison so they could not distort the split. Every male and female percentage in this research is a share of named authors only.

The data was then cut four ways: by version (A, B, C), by platform, by sector, and, for LinkedIn specifically, by content type (pulse article versus post). Representation (share of authors who are women) was also separated from prominence (how high those authors appear in the answer), so that one number could not hide the other.

A measurement note: AI answers vary by question wording, country and platform. Those variables were held constant, so the patterns here reflect that fixed setup. Not every question returned a usable result in every version, so the three versions carry different sample sizes. These are observed patterns in tracked prompts, not guaranteed outcomes. To fully contextualise these findings, it would be necessary to know the gender distribution of LinkedIn content creators publishing in each of the eight sectors studied. A formal request has been submitted to LinkedIn‘s Economic Graph team for this aggregate, anonymised data. Until that data is available, it is not possible to determine with certainty whether the AI visibility gap reflects an underlying publishing gap, amplifies one that already exists, or creates one where there is none.

Key Findings: Seven Results That Reframe the Visibility Problem

The data from Azahara Corrales (AI Governance Strategist) in collaboration with OtterlyAI achieved Women represent only 20% of named LinkedIn authors surfaced by AI Search in neutral (unprompted) expert-discovery queries, a figure that sits at the centre of seven interlocking findings.

Finding 1: LinkedIn Surfaces a Man Four Times Out of Five by Default

LinkedIn is where professionals publish their expertise, and it is the source AI Search leans on most for named people. In the neutral version of the experiment, LinkedIn supplied 38% of all citations that pointed to an identifiable author, more than YouTube, Forbes and Medium combined. Yet when no gender is specified, only 20% of those named LinkedIn authors are women.

The format picture makes it worse. AI Search treats the long-form LinkedIn pulse article as a reference-grade source and cites it far more than short posts. In a companion analysis of more than two million LinkedIn citations, pulse articles took roughly 72% of LinkedIn‘s AI Search citations, the LinkedIn AI Search Citations Study. And it is the format where women are least visible: in my data, women were 27.7% of named authors on LinkedIn posts but just 14.3% on pulse articles.

So the gap is not only about who publishes. It is about who publishes in the format AI Search rewards. Women hold a larger share of the short-post format that AI undervalues, and a smaller share of the long-form articles it cites most. The two effects compound into a visibility deficit that is structural, not incidental.

Finding 2: The Neutral Default Behaves Like a Request for Men

The single-platform shares show the gap. Asking the same question three ways shows that the gap is a property of the default itself, not of the topics chosen. Across every named author AI Search returned, the female share came out like this:

Version

Female share of named authors

A, neutral

24.6%

B, women specified

68.8%

C, men specified

12.9%

Read those three numbers together. A genuinely neutral default would land somewhere between the women-specified and men-specified results. Instead the unprompted answer sits close to the men-specified one, 24.6% against 12.9%, and a long way from the women-specified 68.8%. Asking a question without specifying gender does not produce a balanced list. It produces something close to asking for men.

The flip in Version B confirms the expertise is there. The moment the question says “women”, female share of named authors nearly triples. These are not obscure names scraped from nowhere. They are established practitioners who simply do not surface unless the user thinks to ask for them.

This is what the research calls the prompt tax. Women’s expertise is available inside these systems, but reaching it requires an extra step, specifying gender, that someone searching neutrally never thinks to take. The tax is invisible in everyday use, which is exactly what makes it consequential. The person building a speaker lineup, a media shortlist or a research panel from a neutral AI query does not know they are paying it. They receive a list that looks complete, and they act on it. On LinkedIn specifically, the neutral default is even lower than the cross-source average, at 1 in 5.

Finding 3: Across Publishing Platforms, Women Lead Only on Substack

LinkedIn is the largest source AI Search cites for named experts, but it is not the only one. When the neutral default was compared across the main publishing platforms AI pulled from, the same male skew showed up almost everywhere.

Platform

Female share of named authors, neutral

Substack

79.2%

Instagram

48.9%

YouTube

28.2%

Forbes

20.6%

LinkedIn

20.0%

Medium

9.0%

The written platforms AI Search treats as most authoritative for professional expertise, LinkedIn, Forbes and Medium, are exactly where women are least present. Medium sat lowest at 9%, a striking figure for a platform that positions itself as an independent voice for practitioners and thinkers. The implication is that even platforms designed to democratise publishing have not translated that promise into equal representation in AI search results.

The Substack finding deserves more than a row in a table. It is the only platform where women’s authorship clearly leads, and by a wide margin: 79% of named Substack authors surfaced by a neutral question were women. The pattern held when the data was pooled across all three versions, at roughly 72%. This is a genuine reversal of what every professional platform shows.

Here is the catch. Substack is cited far less by AI Search than LinkedIn or Forbes. Women have built their strongest independent publishing footprint on the platform AI currently leans on least, and the smallest footprint on the long-form professional formats it cites most. That gap, between where women are publishing and where AI Search is looking, is itself the problem. It is also an opportunity. As AI Search indexing evolves and Substack’s weight grows, the women who have built audiences there may find their visibility catching up to their actual influence. For now, the lesson is blunt: a strong Substack following does not yet translate into being surfaced as an expert by AI.

Finding 4: Platform Behaviour Varies, and the Spread Matters

The platforms do not behave the same way, and the differences matter for anyone choosing a tool for research or talent discovery.

Platform

Neutral

Women specified

Men specified

ChatGPT

34.4%

85.5%

18.0%

Perplexity

27.1%

63.5%

10.7%

Google AI Overviews

24.9%

68.8%

12.5%

Google AI Mode

23.2%

68.0%

10.4%

Copilot

16.9%

58.1%

23.0%

Gemini

15.6%

57.7%

12.5%

ChatGPT, the most widely used AI Search tool in the world, returned the highest neutral female share of any platform at 34.4%, and the largest lift when asked for women, up to 85.5%. That is an encouraging finding, and it suggests the default skew is not a fixed property of the technology. Where a platform chooses to surface a broader range of sources, it can.

The weaker performers on the neutral default were Gemini and Copilot, both sitting below 17% female before any gender was specified. Every platform still defaulted male. None broke the pattern. But the spread, from 15.6% to 34.4%, shows that platform design is a variable, not a constant. That distinction matters enormously for governance: if one platform can achieve a higher neutral female share, then the others are not constrained by the data, they are constrained by choices that can be revisited.

Finding 5: The Gap Is Smallest Where the Field Talks About It

The default skew is not uniform across sectors. In the neutral version, female share of named authors by sector ranged widely:

Sector

Female share, neutral

AI and Technology Governance

45.4%

Medicine and Science

36.0%

Personal Development and Coaching

32.7%

Media and Marketing

19.1%

Leadership and Business

18.4%

Finance and Investment

7.7%

Technology (broad)

5.8%

AI and Technology Governance stands out as the one sector approaching balance without prompting. It is also the sector that most actively discusses representation and bias, and that appears to translate into more women publishing visible, citable content on the topic. The 45% baseline reflects a community that has argued for representation, published about it and built networks around it. Deliberate, sustained publishing on a topic creates the citation footprint AI Search rewards.

The fields sitting near 6% female are not there because women lack expertise. They are there because the publishing culture in those fields has not made that shift yet. Finance and Investment (8%) and broad Technology (6%) were the most male-default sectors tested. The gap closes where communities decide to close it, and the AI and Technology Governance sector shows that the decision produces measurable results within a timeframe that matters.

Finding 6: This Is a Retrieval Problem, Not a Merit Problem

When AI Search does surface a woman’s content, it treats it almost the same as a man’s. In the neutral version, women appeared at a very similar position in the answer to men, slightly lower on average but close. The system is not ranking women down once it finds them. It is not finding them in the first pass.

That distinction is the most important one in this study, and it is the most hopeful. Retrieval problems have practical solutions. A merit problem would mean women’s content underperforms once it appears, and it does not. In the macro study, women’s cited content earned slightly more citations per URL than men’s, about 4% more on average. Once a woman’s article is surfaced, it is not cited less. It is found less. The work is getting surfaced at all, and that is a problem GEO can act on.

Finding 7: A Five-Month Macro Study Points the Same Way

This experiment is a snapshot of expert-discovery questions. It sits inside a larger picture. The companion LinkedIn AI Search Citations Study tracked the gender of authors behind LinkedIn content cited by AI Search across five months and every industry, not just questions that ask for experts. The two studies agree, which matters: a controlled prompt test and a broad observational study landing in the same place is stronger evidence than either alone.

In the macro study, women are 23.5% of the named LinkedIn authors AI Search cites across all topics. In this experiment, the neutral default puts women at 20.0% of named LinkedIn authors. Both sit in the same 1 in 4 to 1 in 5 band.

LinkedIn named authors

Macro study, all industries

This experiment, expert-discovery prompts

Female share

23.5%

20.0%

Male share

76.5%

80.0%

The small gap between the two is the interesting part. When the question is specifically who the experts are, women’s share dips below their share of LinkedIn citations in general. Expert framing does not lift women’s visibility. It nudges it slightly lower.

The macro study also shows how steady the gap is across engines. Across Perplexity, Google AI Overviews, ChatGPT, Google AI Mode and Copilot, the female share lands within a single percentage point of 23 to 24%. This is not one engine pulling the average down. It is a systemic default that every major platform reproduces at almost exactly the same level.

AI Search Platform

Female share of person citations

Perplexity

23.3%

Google AI Overviews

24.2%

ChatGPT

23.1%

Google AI Mode

23.9%

Microsoft Copilot

23.0%

Gemini

43.3%

Overall

23.5%

The consistency is the real finding here. Across Perplexity, Google AI Overviews, ChatGPT, Google AI Mode and Copilot, the female share lands within a single percentage point of 23 to 24%. Gemini reads at 43.3%, but on a negligible share of citations, so that should not be treated as a real platform difference. Broken out in full, the male, female and non-binary composition of person citations on each platform looks like this:

Gender

Perplexity

Google AI Overviews

Chat-
GPT

Google AI Mode

Microsoft Copilot

Gemini

Overall

Male

76.7%

75.7%

76.9%

75.9%

77.0%

56.7%

76.4%

Female

23.3%

24.2%

23.1%

23.9%

23.0%

43.3%

23.5%

Non-binary

0.0%

0.1%

0.1%

0.1%

0.1%

0.0%

0.1%

Non-binary authors are a rounding error on every platform, under one in a thousand citations. The story is a male default that holds almost identically wherever AI Search looks.

What This Means for GEO and AI Governance

Generative Engine Optimisation is about being surfaced as a source inside AI answers, not just ranking on a results page. When AI Search builds an answer about experts or voices to follow, it selects from the content it can find, parse and trust. If that selection skews male by default, then any team using AI Search to build a speaker list, a media list, an analyst shortlist or a hiring pipeline inherits that skew without seeing it. That is not a neutral outcome. It is a governance failure dressed as a search result.

For women professionals, this turns AI visibility into a concrete publishing problem, and a solvable one. The highest-leverage action is to publish long-form LinkedIn articles consistently, with a clear byline, declared expertise on the profile, and topic language that matches how people actually search. Short posts build community but rarely build AI citability. The format AI rewards is the one where women are currently least represented, which means it is also the biggest opportunity.

Working alongside professionals navigating AI visibility, the pattern that emerges is consistent: the gap between where women publish and where AI Search looks is not a talent deficit. It is a structural mismatch between publishing habits and citation mechanics. Corrales’s background, with a degree in journalism and a Master’s in international marketing, alongside her speaking career that has evolved from marketing-focused talks to governance, legislation and responsible AI adoption at major conferences including Brighton SEO, gives this research a practitioner-researcher perspective that is rare in this field. She is not observing from outside. She is inside the problem.

With the EU AI Act now in force, three implications follow for anyone writing or applying governance policy. First, measurement has to be explicit: representation in AI answers should be audited per sector and per platform, because a blended average hides a gap that is not uniform. Second, retrieval deserves scrutiny, not just training data; the bias here lives in which sources get surfaced for a neutral query, which is a measurable and addressable layer. Third, any organisation using AI Search to identify experts for panels, hiring or funding is making representation decisions through a tool that defaults male, often without knowing it. Naming that default is the first governance step. Designing prompts and review processes that correct for it is the second.

The platform differences add a fourth point. If tools vary this much in how they surface women, organisations choosing an AI Search tool for internal research or talent discovery should be asking vendors for representation audits by sector. The data to produce them exists. The question is whether it will be demanded.

This is also a question about who builds these tools. A team that is predominantly male, working within a predominantly male professional network, is less likely to notice that their tool fails to surface women, because their own searches are already returning results that look complete to them. Diverse teams are not just an equity goal. They are a quality control mechanism.

Key Takeaways: What to Do With This Research

The findings from this study are initial. The direction has held steady throughout collection. And the response is practical, not abstract.

If You Are a Woman Building Your Professional Visibility

The problem is retrieval, not merit. Write long-form LinkedIn articles on your area of expertise. Use the language people search for, not only the language your field uses internally. Declare your credentials and topic focus clearly on your profile. Do it consistently. You are not fighting the system. You are giving it what it needs to find you.

If You Use AI Search to Find People

Any list you build from an unprompted query starts from a male-skewed default. Treat it as one input, not the answer. Run gendered variants of the same question and compare. The expertise you are missing surfaces immediately when you ask for it. That step takes thirty seconds. It changes who ends up on your shortlist.

If You Work in GEO or AI Governance

Default retrieval is not neutral, and treating AI Search as a single blended channel hides the gap. Segment by sector and by platform. The same query returns very different representation depending on the field it sits in and the engine answering it. The data to audit this exists. The question is whether the organisations commissioning AI-assisted research will demand it from their vendors, and whether governance frameworks will require it as a standard disclosure.

AI Search does not rank women lower. It leaves them out of the default and surfaces them mainly when asked. On LinkedIn, the source it relies on most for named experts, women are just 1 in 5 of the people it surfaces by default, and only 1 in 7 on the long-form articles it cites most. The expertise is present and, once surfaced, gets cited at nearly the same prominence. The gap is one of retrieval. That makes it a problem GEO can act on, and a responsibility governance cannot ignore.

The gap closes where people decide to close it, through consistent, citable, long-form content that AI can find and trust, and through organisations that stop accepting the neutral default as a complete answer.

 

ABOUT THE AUTHOR

Azahara Corrales

Azahara Corrales is an AI governance strategist, speaker and author. Her work focuses on responsible AI and women's leadership in AI, and she is the creator of the MATRIZ framework for AI governance. Recognised as Digital Marketer of the Year and Women to Watch 2023 by the Digital Women Awards, and shortlisted as AI Specialist Finalist, she has spoken at Brighton SEO four times, as well as at AfricaTech and DES Málaga. Her forthcoming book, Nobody Told Me This Was For Me: Why Women Should Be the Next Leaders of AI and How to Get Started, develops the research and arguments presented here. This study was produced in collaboration with Rick Tousseyn from OtterlyAI, which provided the measurement layer for tracking how experts appear in AI answers across ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Gemini and Microsoft Copilot.

RESEARCH PARTNER

About OtterlyAI

This research was conducted in collaboration with Rick Tousseyn, Pam Tarrayo and OtterlyAI, an AI search monitoring and Generative Engine Optimisation platform. OtterlyAI provided the measurement infrastructure for this study, tracking how content and experts appear in AI-generated answers across ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Gemini and Microsoft Copilot.