Stop using Media Bias/Fact Check in research


1 Introduction↩︎

Social media has ushered in the so-called post-truth era, and, in response, the burgeoning field of misinformation science. Like all scientific endeavors, as part of its inquiry, the field must count, classify, and measure the properties of misinformation, the infrastructure for which is expanding. We examine a popular piece of that infrastructure, Media Bias/Fact Check (MBFC).

MBFC is widely accepted in scientific practice [1][8], recommended as a tool by university libraries [9][11], bundled into composite datasets [12][15], and used in commercial products (such as the popular consumer news aggregation site ground.news [16]). [8] aptly summarize this consensus in a footnote, writing “MBFC is currently the most comprehensive media bias resource on the internet,” an almost identical statement to that on MBFC’s own homepage [17].

MBFC rates the {bias} of media institutions along a left-to-right political spectrum. Despite this framework’s popularity, it is not obvious that the political bias of institutions across the globe can be meaningfully distilled to a single point in one dimension. This concept of bias contains many complex, overlapping assumptions, both in what it includes and ignores. For example, MBFC’s rubric focuses only on the content published by each institution, disregarding its governance structure and finances.

The roots of this definition of {bias} can be traced to the Progressive era reformers’ critique of the press. From there, it was iteratively transformed, passing through McCarthyism and the civil rights movement, until becoming foundational to the coalescing conservative movement, which undertook a sustained public relations campaign to discredit the media for its supposed liberal bias [18][20].

In the 1940s, media critics like George Seldes and Morris L. Ernst warned of the increasing concentration of ownership in newspapers, with both putting forth critiques stemming from structural causes [18]. Seldes’s In Fact argued that corporate news, because of its ownership structure, is systematically biased in favor of capital. Ernst, a prominent ACLU lawyer, argued that the high cost of producing media threatened access to freedom of speech, and, therefore, a meaningful implementation of it. As McCarthyism took hold, Seldes was successfully sidelined, but, taking a path emblematic of our story, Ernst’s own anti-communism led him on a more complicated path, along which he brought his structural lens, and which eventually took him to Accuracy in Media, which we discuss shortly [18].

During the 1950s, the nascent conservative coalition was in the process of assembling itself [18], [19]. As the US civil rights movement progressed, local papers in the Southern US routinely downplayed the violence that the state directed at civil rights protesters. By contrast, the incipient national media, like the New York Times, printed sympathetic coverage, outraging white Southerners, who accused the national media of bias and unfairness [18], [19].

Meanwhile, the rise of the John Birch Society and the Barry Goldwater campaign of 1964 showed an insurgent faction of an anti-communist, proto-conservative movement [18], [19], [21]. The national media of the time was critical of Goldwater, with the media calling Goldwater an extremist, and Goldwater accusing the media of being part of a biased, radical elite [19]. By the mid-1960s, the overlapping white Southern, Goldwaterite, and Bircher factions had, each for their own reasons, developed critical dispositions towards the media.

As the civil rights movement made it increasingly taboo to explicitly support white supremacy, this coalition could, in their attacks on the media, fall back to the more abstract accusation of liberal bias, which found resonance in all camps [19]. Archsegregationist Alabama Governor George Wallace, for example, frequently attacked “the liberal, left-wing press” [19]. As [18] writes:

[T]he media activism of the John Birch Society [...] demonstrates how the “liberal media” claim emerged less as a top-down movement strategy than as a bottom-up interpretation of the various headwinds facing modern conservatism as it navigated rapidly shifting conditions regarding the maintenance of white supremacy, particularly in the US South.

Parallel to this bottom-up work, there was a developing, top-down-funded intellectual movement to complement it. Oil tycoon H.L. Hunt’s Facts Forum, an anti-communist, proto-conservative organization and publication, began an intellectual project not just to critique the media directly, but to foster a critical posture towards the media among a coalescing conservative community. Among its many tactics, Facts Forum took advantage of the Fairness Doctrine to receive airtime for programming that supposedly presented both sides of issues, but whose presentation and framing clearly disadvantaged the left [18].

William F. Buckley, a former Facts Forum reporter, would go on to found the most influential conservative media outlet, the National Review [18]. Buckley pioneered...

... a conservative form of respectability politics [...] Conservatism’s association with the interests of big business, and with carrying water for fascists abroad in the run up to World War II, inhibited the movement’s growth beyond a minority status [...] Buckley and those in his orbit successfully reframed modern conservatism as a principled and well-disciplined movement beset by a liberal “establishment” (including the media) and a right-wing “fringe’,” exemplified by [...] groups like the John Birch Society [18].

This nascent coalition had its tensions, but the rising conservative intelligentsia, along with the various aforementioned strands of the burgeoning conservative movement, found common cause in hostility to the press.

In 1969, with the “liberal media” critique established in the conservative movement, Reed Irvine founded Accuracy in Media (AIM), which “helped make media bias an explicit conservative movement cause” [18]. Originally, AIM spun off from the Council Against Communist Aggression (CACA), through which it attracted Ernst and other liberal and progressive media reformers who brought with them the remnants of the Progressive era’s structural media criticism. Despite its CACA origins and initial putative bipartisanship, AIM’s criticism of the media quickly found a constituency in the conservative movement. AIM became increasingly aligned with this movement as President Nixon and Vice President Agnew regularly attacked the media for its liberal bias, which, in their telling, interfered with the media’s duty to tell the truth [18].

AIM marks the beginning of the institutionalization of a critique that both rests on and sustains the assertion that the media is liberal [18]. Such research is often funded by big business, which, as a rule, prefers not to be investigated, and is therefore generally hostile to an adversarial press [20]. By the time that the Center for Media and Public Affairs (CMPA) began publishing its Media Monitor in 1985, this critique was fully institutionalized. CMPA’s research aimed to “demonstrate the liberal bias and anti-business propensities of the mass media” [22]. Their work often employs dubious methodologies designed to reach the predetermined outcome:

[T]he main analytical technique used by the Center—the counting of “thematic messages”—is extremely dubious, eliminating all messages that fail to make an explicit statement of opinion. Since sources who accept the status quo don’t need to explicitly state an opinion, this technique often produces highly distorted findings. For example, the CMPA report on Gulf War coverage found that “nearly three out of five sources (59 percent) criticized U.S. government policies during the Gulf War.” This improbable result comes from throwing out 5,666 of 5,915 messages, and looking only at what the remaining 249 said about U.S. policy. [23]

CMPA, AIM, and other think tanks continue this work to the present day [20], [23].

In the 1990s, Rupert Murdoch founded his network, Fox News, explicitly to combat the supposed liberal bias of the media, as did other conservative media moguls, like Roger Ailes and Conrad Black, all of whom built influential media empires [20]. With Murdoch, we see the concept of the liberal media become fully mainstream. This mainstream acceptance provides inherent credibility to the founding of explicitly rightist media organizations aiming to balance liberal media, even as rightist media organizations increasingly dominate the media industry [24].

0.48

MBFC’s guides for interpreting (left) and (right). For the latter, we remove rows 3–7 for brevity.
Score Rating Description
0 Very High Consistently factual, uses credible information, no failed fact checks.
0.1–1.9 High High factual, minor sourcing issues, reasonable fact check record.
2.0–4.4 Mostly Factual Generally reliable but may have occasional fact-check failures, transparency, and sourcing issues.
4.5–6.4 Mixed Reliability varies; multiple fact-check failures, poor sourcing, lack of transparency, one-sidedness.
6.5–8.4 Low Often unreliable; frequent fact-check failures and significant issues with sourcing, transparency, propaganda, conspiracies, and pseudoscience promotion.
8.5–10 Very Low Consistently unreliable, heavily biased, with intentional misinformation likely.

0.48

MBFC’s guides for interpreting (left) and (right). For the latter, we remove rows 3–7 for brevity.
Score Description
0 Perfect balance, presenting all sides equally with no discernible bias or use of emotional language.
1 Almost perfectly balanced, with very minor favoritism and minimal use of subtle emotional cues, but all perspectives fairly represented.
2 Minor bias, slightly favoring one side with occasional use of emotionally suggestive terms, while still maintaining reasonable balance and representation of opposing views.
8 Heavy bias, consistently favoring one side with little effort to present alternative viewpoints and pervasive use of emotional language that borders on propaganda.
9 Strong bias, rarely including alternative perspectives, using extreme emotional framing or language that seeks to manipulate reader perception almost entirely in favor of one viewpoint.
10 Extreme bias or propaganda, exclusively presenting one side with no balance or acknowledgment of opposing views, often employing inflammatory, divisive, or manipulative language to an extreme degree.

As we will show, MBFC’S methods and conclusions represent and entrench this intellectual lineage. That MBFC’s data is prevalent throughout the misinformation literature is a testament to the success of the aforementioned attacks on the press, and to the failure of much misinformation scholarship to understand the political context of its work.

To that end, we must make clarifying note about our intention: The present study is highly critical of MBFC, but this criticism is aimed at the scholarship that relies on it. Good scholarship does not necessarily come from the academy, but it is always the result of mentorship, collaboration, critique, and community. The prevalence and usage of MBFC in peer-reviewed literature shows that these processes have broken down in the very institutions designed to preserve and expand them. What could have been a productive interdisciplinary collaboration in a time of waning public trust in science has instead become an ironic case study.

To explore MBFC’s design and impact, we first outline MBFC’s methodology and data (Section 2). We then quantify the presence of MBFC in the misinformation literature (Section 3), creating a database of papers with which we explore why authors use MBFC (Section 4). In Section 5, we argue that MBFC’s widespread usage in the literature is due not to its accuracy, but to providing specious data to researchers who are themselves downstream of the same processes that shaped MBFC, and who therefore conflate the familiarity of its conclusions with accuracy. We also argue that this resonance allows MBFC to bypass critical examination, as demonstrated by the widespread mischaracterization of MBFC throughout the literature, with published descriptions of MBFC often contradicting MBFC’s own website.

2 MBFC’s Methodology and Data↩︎

MBFC has a methodology page [25] that outlines the various rubrics used to generate its data. The page has changed considerably since it was initially published. For example, MBFC has, in the past, implemented a voting system in which users can vote on the {bias} of different sources, which it would then display and even take into account in producing their official ratings [26]. Until 2025, however, the methodology page was relatively scant, though in 2024 we start to see the beginnings of the rubrics we will discuss, e.g., references to {economic system} [27].

In 2025, MBFC redesigned the rating systems, making it much more highly specified, with rubrics and scoring. This basic structure is the one in use as of this writing [25], [28].

When discussing MBFC’s methodology, we will discuss only the most up to date version, as of this writing. However, it is important to note that most of the papers found in Section 3 predate the present rubric, when MBFC’s lack of methodological rigor was impossible to dispute.

We will discuss MBFC’s measurements of {factual reporting}, {bias}, {sec:credibility}, and {freedom}. The first three are available through a single API endpoint, which returns a CSV file rating \(9,365\) sources. MBFC provides these scores for a diverse set of outlets, but the typical example is a media organization, such as CNN or Fox News, though one also finds unions and think tanks. The {freedom} score is for countries, not sources, and it comes directly from the MBFC site and is not available through the site’s API. For important context, we discuss each of these concepts’ design and score paradigms.

0.48

MBFC’s scoring rubric (left) and interpretation criteria (right) for . Note that, in the top left, ’s lowest score () is also omitted in the original. Note also the nested logic on the right, which overrides the parent rubric.
Category Rating Points
Factual Reporting Very High 4
High 3
Mostly Factual 2
Mixed 1
Low 0
Bias Least Biased / Pro-Science 3
Right-Center or Left-Center 2
Left or Right 1
Questionable/Conspiracy/
Pseudoscience
0
Traffic/ Longevity High Traffic 2
Medium Traffic 1
Minimal Traffic 0
Bonus: \(\geq 10\) years existence +1
Press Freedom Limited Freedom -1
Total Oppression -2

0.48

MBFC’s scoring rubric (left) and interpretation criteria (right) for . Note that, in the top left, ’s lowest score () is also omitted in the original. Note also the nested logic on the right, which overrides the parent rubric.
Credibility Level Criteria
High Credibility A score of 6 or above.
Medium Credibility A score between 3–5 points. Additionally, in accordance with the MBFC scoring system, a “Mostly Factual” rating with a score between 3.6 and 4.5 automatically results in a Medium Credibility classification, regardless of the overall tally in other categories. This reflects the critical importance of factual accuracy in determining credibility.
Low Credibility A score of 0–2 points. Sources rated as “Questionable,” “Conspiracy,” or “Pseudoscience” are automatically classified as Low Credibility.

2.1 Factual reporting↩︎

According to MBFC’s methodology page, the {factual reporting} score is based on a 10-point scale, created using the following rubric: “Failed Fact Checks (40%), Sourcing (25%), Transparency (25%), and One-Sidedness/Omission (10%).” In what is a theme throughout this section, there is no way to access the original scale. The API only provides access to the ratings presented in Table ¿tbl:tab:factual-reporting?, from \(\boldsymbol{\langle}\) very low \(\boldsymbol{\rangle}\) to \(\boldsymbol{\langle}\) very high \(\boldsymbol{\rangle}\), and the transformation is irregularly binned. The bin size for \(\boldsymbol{\langle}\) high \(\boldsymbol{\rangle}\), for example, is from 1–1.9, but \(\boldsymbol{\langle}\) mostly factual \(\boldsymbol{\rangle}\) is from 2.0–4.4. No explanation is given for this transformation.

Table ¿tbl:tab:factual-reporting? shows the {one-sidedness} portion of the {factual reporting} score. Despite MBFC reporting an ostensibly separate {bias} score (which we discuss in Section 2.4), 10% of the {factual reporting} score contains what is conceptually difficult to distinguish from {bias}. We discuss this in Section 5.

Finally, a small note on notation: When plotting {factuality}, we use a 0–5 scale, where 0 is \(\boldsymbol{\langle}\) very low \(\boldsymbol{\rangle}\) and 5 is \(\boldsymbol{\langle}\) very high \(\boldsymbol{\rangle}\). This is an inverted scale from MBFC’s internal one, in which smaller numbers correspond to higher factuality, but, since we do not have access to those numbers anyway, we have chosen the more intuitive alternative.

2.2 Credibility↩︎

The methodology for calculating {sec:credibility} is shown in Table ¿tbl:tab:credibility?. It is a composite of {factual reporting}, {bias}, and {freedom}, along with the site’s traffic, for which they use an estimate of page views. Though this process generates a numerical score, we are given access to a binned version with uneven bin sizes described in Table ¿tbl:tab:credibility?. Note that this table contains nested logic, with 2 of the branches resulting in an override to what is described in the parent rubric.

2.3 Freedom↩︎

MBFC assigns each country a {freedom} score. The methodology is as follows [29]:

We calculate the overall freedom levels of countries by averaging the scores from Reporters Without Borders (RSF) and Freedom House’s yearly Freedom in the World Index [...] If a country has not been rated by either source, we use other sources such as Unesco.org, Statista.org, BBC Country Profiles, or other credible sources to estimate the level of freedom.

The methodology uses a point score system from 0–100. However, when downloading the data, the original 100-point score is unavailable. Instead, it is grouped into five categorical labels: Total oppression (0–24), Limited freedom (25-49), Moderate freedom (50-69), Mostly free (70-89), and Excellent freedom (90-100).

The uneven bucket size is unexplained on the methodology page [29], and the raw score is unavailable through the API. Thus, available data is on a five-point scale of uneven bin size in the original rating space. For our analysis, we label these bins 0–4, with 0 being “Total Oppression” and 4 being “Excellent Freedom.”

We consider the freedom scores for “countries with which the United States has strained or hostile relations”, including: Afghanistan, Belarus, China, Cuba, Eritrea, Iran, Myanmar, Nicaragua, North Korea, Russia, South Sudan, Syria, and Venezuela [30]. MBFC rates every one of these countries’ freedom as “total oppression” except for one, South Sudan, which reaches “limited freedom”. Similarly, NATO countries have a mean score of \(3.1\) whereas non-NATO have a mean score of \(2.0\).

2.4 Bias↩︎

Figure 1 shows counts of sources in MBFC’s dataset grouped by their {bias} ratings. This scale presents problems of interpretation. The left-to-right values are, as one might expect, ordinal values that can be arranged along a one-dimensional spectrum, but it is difficult to explain or understand why \(\boldsymbol{\langle}\) pro-science \(\boldsymbol{\rangle}\), \(\boldsymbol{\langle}\) questionable \(\boldsymbol{\rangle}\), \(\boldsymbol{\langle}\) satire \(\boldsymbol{\rangle}\), and \(\boldsymbol{\langle}\) conspiracy-pseudoscience \(\boldsymbol{\rangle}\), categorical values outside that spectrum, appear as potential {bias} values.

Figure 2 also compares {bias} to {factual reporting}. Taken together with Figure 1, researchers might find a familiar idea, that of the “liberal media” from Section 1, which we discuss further in Sections 5 and 6.

Figure 1: Distribution of {bias} categories in MBFC. Categorical values outside the left/right spectrum are separated out to simplify interpretation. Note that the two are exclusive, e.g., a source cannot be both \boldsymbol{\langle} least biased \boldsymbol{\rangle} and \boldsymbol{\langle} pro-science \boldsymbol{\rangle}, or \boldsymbol{\langle} center-right \boldsymbol{\rangle} and \boldsymbol{\langle} questionable \boldsymbol{\rangle}.
Figure 2: {bias} vs. {factual reporting}. The categorical values outside the left/right spectrum are to the right, represented in green, whereas the left/right values are to the left in gold.

Further complicating the interpretation, MBFC’s methodology page (see Table ¿tbl:tab:bias-schema?) does not include these extra categories. It shows instead a {bias} score that goes from \(-10\) to \(+10\), but which is then binned into a 7-point scale. The actual data exposed to the user, however, contains a 5-point scale (once the values mentioned above are removed). The original \(-10\) to \(+10\) scale seems internal to MBFC and is never exposed to users. Within the explanation of the 7-point scale, however, the second and second-to-last items contain parentheticals which we believe are intended to explain the 5-point scale.

To summarize, there are three different scales used: a twenty-point scale, a seven-point scale, and a five-point scale. The twenty-point scale seems to be the original, internal scale, which gets coarse grained into seven- and five-point scales. The seven-point scale seems to exist only in the explanation for the bias score. The five-point scale is potentially explained in parentheticals in that same methodology section, and it is the data with which MBFC’s users directly interact, with the caveat that it also contains the four extra values that do not sit anywhere on the spectrum.

For some clarity on these four extra values, we can turn to the {sec:credibility} rubric (discussed in Section 2.2 and in Table ¿tbl:tab:credibility?). Recall that {sec:credibility} is a composite score that includes {bias}. A source is more credible if it is either \(\boldsymbol{\langle}\) least biased \(\boldsymbol{\rangle}\) or \(\boldsymbol{\langle}\) pro-science \(\boldsymbol{\rangle}\), for which it earns 3 points. If it is \(\boldsymbol{\langle}\) right-center \(\boldsymbol{\rangle}\) or \(\boldsymbol{\langle}\) left-center \(\boldsymbol{\rangle}\), it gets 2 points. For \(\boldsymbol{\langle}\) left \(\boldsymbol{\rangle}\) or \(\boldsymbol{\langle}\) right \(\boldsymbol{\rangle}\), 1 point; and for \(\boldsymbol{\langle}\) questionable/conspiracy/pseudoscience \(\boldsymbol{\rangle}\), which we assume is a composite of \(\boldsymbol{\langle}\) questionable \(\boldsymbol{\rangle}\) and \(\boldsymbol{\langle}\) conspiracy-pseudoscience \(\boldsymbol{\rangle}\), it gets 0 points.

To calculate {bias}, MBFC uses the following scoring categories [25]:

The placement of a source on the Left-Right Bias Scale is determined by a weighted composite score derived from four categories: Economic System (35%), Social Progressive Liberalism vs. Traditional Social Conservatism (35%), Straight News Reporting Balance (15%), and Editorial Bias (15%). Scores are on a scale of \(-10\) to \(+10\), and the weighted average determines the overall bias score.

MBFC’s categories and score definitions.
Score Range Bias Category
\(-10\) to \(-8.0\) Extreme Left Bias
\(-7.9\) to \(-5.0\) Left Bias (Far Left at \(-7.0+\) [sic])
\(-4.9\) to \(-2.0\) Left-Center Bias
\(-1.9\) to \(+1.9\) Least Biased
\(+2.0\) to \(+4.9\) Right-Center Bias
\(+5.0\) to \(+7.9\) Right Bias (Far Right at \(+7.0+\))
\(+8.0\) to \(+10\) Extreme Right Bias
MBFC’s categories and score definitions.
Score Description
\(-10\) Communism: Advocates no corporatism, extreme regulation, and full government ownership of industries.
\(-7.5\) Socialism: Supports minimal corporatism, high regulation, and significant government ownership.
\(-5\) Democratic Socialism: Endorses reduced corporatism with strongly regulated capitalism.
\(-2.5\) Regulated Market Economy: Promotes moderate corporatism with balanced regulations.
\(0\) Centrism: Balances regulation and corporate influence without significant bias.
\(2.5\) Moderately Regulated Capitalism: Leans slightly toward corporatism with moderate government intervention.
\(5\) Classical Liberalism: Emphasizes moderate to high corporatism with lower regulations.
\(7.5\) Libertarianism: Advocates low government intervention and high corporate influence.
\(10\) Radical Laissez-Faire Capitalism: Advocates minimal to no regulation, with the economy governed entirely by free-market principles and private enterprise.

We provide two scoring category examples from the full {bias} rubric. First, the {economic system} rubric portion, accounting for the plurality of the score at 35%, can be found in Table [tbl:tab:economic-system]. In it, MBFC labels the farthest left as “Communism,” defined as “[advocating] no corporatism, extreme regulation, and full government ownership of industries.” The farthest right is “Radical laissez-faire capitalism,” which “[a]dvocates minimal to no regulation, with the economy governed entirely by free-market principles and private enterprise.”

Second, the {straight news reporting balance} rubric portion...

[m]easures how well a source reports all sides in its straight news stories, either through story selection or content balance within articles. This covers strictly news reporting and is separate from Editorial/Op-Ed bias.

We discuss these examples further in Section 5.

3 MBFC in the literature↩︎

A systematic analysis of MBFC in the scientific literature is difficult because not all papers that rely on MBFC cite or credit it in the same way. To illustrate this issue, searching citations databases Semantic Scholar [31], OpenAlex [32], OpenCitations [33], and Scopus [34] for “mediabiasfactcheck.com” or “Media Bias/Fact Check” yields relatively few results. For these strings, respectively, Semantic Scholar shows 2 and 21 results, OpenCitations shows 0 and 0, OpenAlex shows 7 and 0, and Scopus shows 1 and 142 results. Meanwhile, Google Scholar returns \(1,610\) results for “Media Bias/Fact Check.”

The discrepancy between Google Scholar and the various metadata searches suggests that many papers that use MBFC as a dataset do not have its citation in their references, at least not in the format the metadata-based citation databases are expecting.

We choose not to use the results from Google Scholar for two reasons. First, its indexing and algorithm are opaque, so much so that we find the total number of results for the same query fluctuates day to day. Second, we find its permissiveness towards web scraping similarly variable. We conclude that these factors would harm the reproducibility of our work.

More importantly, citations represent the official scientific account of how knowledge spreads. One of our main arguments is that MBFC’s methodology is flawed, and that, through its continued use in research, MBFC’s scores become divorced from its methodology such that the provenance is effectively erased. Therefore, we use citations because they are the mechanism by which readers of studies ought to be able to trace such processes.

3.1 Methods↩︎

To estimate the prevalence of MBFC in the scientific literature, we use snowball sampling (see Figure 3). We seed this snowball sampling process with the papers exported from the Scopus search for “Media Bias/Fact Check.” For each seed paper, we create a Paper object in a database. We then fetch every paper that our seed paper cites (“outbound citations”) or that cites our seed paper (“inbound citations”). For the fetched papers, we also create Paper objects, which we link to the seed papers through a Citation object, a directional link between the two papers.

We then attempt to fetch the manuscript text of each fetched paper. If successful, we scan this text for mentions of MBFC. If the paper contains MBFC, we add it to our seed papers and recursively repeat the algorithm, again outlined below for clarity:

  1. For each paper, find both inbound and outbound citations.

  2. For both the resulting inbound and outbound citations, use metadata databases OpenAlex, Crossref, and OpenCitations to find metadata and a link to a PDF version of the paper.

  3. Fetch the PDF, if possible.

  4. Scan the manuscript text of the paper for references to MBFC.

  5. If the paper contains a reference to MBFC, return to step 1.

We then run the sampler to its natural end, i.e., until no new papers are found to re-seed the algorithm.

Figure 3: The snowball sampling process looks at the text of a paper. If it finds a reference to MBFC, it fetches all its references and recursively begins the process anew on each one.

3.2 Results↩︎

We find MBFC present in 3.50% (\(372\)) of the \(10,642\) papers in our sample (or 5.0% of papers if we exclude those for which we could not fetch the manuscript text; see Table 1). We expect that 3.50% is an underestimate of the impact of MBFC in the misinformation literature for the following reasons.

Table 1: Counts of papers with MBFC in the entire database and in the subset of papers containing the term “misinformation” in their title, abstract, or body.
Entire database Papers containing “misinformation”
Papers with text All papers Papers with text All papers
Papers with MBFC 329 (5.01%) 372 (3.50%) 259 (7.38%) 263 (6.49%)
Total papers 6,514 10,642 3,509 4,051

No paper text available: First, for 61% of papers, we are unable to fetch the actual manuscript text, terminating the snowball sample for that branch (even though subsequent connected papers could have been relevant to this analysis). Similarly, detecting MBFC in a paper involves scanning its text, meaning that there are almost certainly papers in our database that use MBFC that we cannot observe. Even in the case that we do find the PDF, scanning its text comes with many limitations, including problems with formatting and inconsistent OCR.

Greediness of snowball algorithm: Further diluting the percentage, the greediness of the methodology biases our database towards highly cited papers. Figure 4 shows the rank distribution of papers in our database. Most of the highly cited papers in our database are outside of the field of misinformation studies, diluting the percentage scores. Conversely, we suspect that there are papers in the literature with low citation counts using MBFC that we did not find.

Bundled sources including MBFC: Our study only takes into account studies that use MBFC directly. We know of at least 4 sources that combine MBFC with other data and some interpretative work to create composite databases [12][15]. As in the case of MBFC, these data sources can be difficult to trace through citations and metadata. A more complete accounting of MBFC’s usage in the literature would require repeating the snowball process on a complete list of secondary sources of MBFC.

Google Scholar: We do not use Google Scholar’s counts directly for the reasons outlined previously. By comparison to our count of \(372\), Google Scholar returns roughly \(4.3\times\) this count at \(1,610\) results for the search query “Media Bias/Fact Check.”

Figure 4: Rank by citation count, with the most cited paper as rank =1. Exponential decay is fast (\alpha=1.93), reflecting our snowball algorithm’s preference for highly-cited papers. Titles of papers whose manuscript text contains MBFC are in gold. Those without are in green.
Table 2: Most cited papers using MBFC in our database with citations \(\ge50\) according to [33]. For additional context, we also include citation counts from Google Scholar [35].
Google Scholar Citations Open Citations Citations Title
\(3395\) \(1320\) The Covid-19 Social Media Infodemic
\(3824\) \(845\) The Echo Chamber Effect On Social Media
\(2092\) \(630\) A Survey Of Fake News
\(1487\) \(570\) Influence Of Fake News In Twitter During The 2016 Us Presidential Election
\(986\) \(367\) Beyond News Contents
\(775\) \(361\) Assessing The Risks Of ‘Infodemics’ In Response To Covid-19 Epidemics
\(787\) \(249\) “Fake News” Is Not Simply False Information: A Concept Explication And Taxonomy Of Online Content
\(976\) \(241\) Auditing Radicalization Pathways On Youtube
\(419\) \(221\) Covid-19 Vaccine Hesitancy On Social Media: Building A Public Twitter Data Set Of Antivaccine Content, Vaccine Misinformation, And Conspiracies
\(581\) \(203\) Fighting An Infodemic: Covid-19 Fake News Dataset
\(613\) \(184\) Analyzing The Digital Traces Of Political Manipulation: The 2016 Russian Interference Twitter Campaign
\(337\) \(153\) The Stealth Media? Groups And Targets Behind Divisive Issue Campaigns On Facebook
\(456\) \(151\) FANG
\(372\) \(150\) Fake News Detection Based On News Content And Social Contexts: A Transformer-Based Approach
\(325\) \(127\) Sentiment Analysis For Fake News Detection
\(284\) \(125\) The Covid-19 Infodemic: Twitter Versus Facebook
\(444\) \(121\) Examining The Alternative Media Ecosystem Through The Production Of Alternative Narratives Of Mass Shooting Events On Twitter
\(377\) \(114\) You Are Fake News: Political Bias In Perceptions Of Fake News
\(312\) \(111\) Recovery
\(215\) \(99\) False News On Social Media
\(207\) \(83\) Evaluating Deep Learning Approaches For Covid19 Fake News Detection
\(365\) \(70\) Combating Disinformation In A Social Media Age
\(143\) \(67\) A Hybrid Linguistic And Knowledge-Based Analysis Approach For Fake News Detection On Social Media
\(264\) \(67\) Exploring The Role Of Visual Content In Fake News Detection
\(167\) \(64\) Search Bias Quantification: Investigating Political Bias In Social Media And Web Search
\(148\) \(64\) Who Falls For Online Political Manipulation?
\(145\) \(63\) Red Bots Do It Better:Comparative Analysis Of Social Bot Partisan Behavior
\(200\) \(60\) Nela-Gt-2018: A Large Multi-Labelled News Dataset For The Study Of Misinformation In News Articles
\(153\) \(59\) Overview Of Constraint 2021 Shared Tasks: Detecting English Covid-19 Fake News And Hindi Hostile Posts
\(160\) \(57\) Credibility-Based Fake News Detection
\(186\) \(56\) Characterizing The 2016 Russian Ira Influence Campaign
\(101\) \(50\) Understanding High- And Low-Quality Url Sharing On Covid-19 Twitter Streams
\(96\) \(50\) Health Misinformation Detection In The Social Web: An Overview And A Data Science Approach

In an attempt to better estimate MBFC’s impact in the relevant literature, instead of looking at all papers uncovered by the sampling methodology described above, we include only those that additionally contain the word “misinformation” somewhere in the title, abstract, or manuscript text. In this subset, 4,051 papers of the original 10,642 remain, of which 263 (6.49%) reference MBFC. If we further filter those to papers for which we have manuscript texts, we find 3,509, of which 259 (7.38%) reference MBFC. Note that \(372 - 259 = 113\) papers contain MBFC but are filtered out of this subset.

Table 2 contains the titles of papers in which we detected MBFC that have 50 or more citations, according Open Citations [33]. Of the 32 papers listed, 8 are developing methodologies for detecting misinformation, meaning that MBFC is used in the literature as a ground truth for machine learning models. If these models are subsequently used in other studies, or perhaps in industry, there is a risk of encoding MBFC’s methodology and conclusions into future work. Due to aforementioned constraints in citation data, the original methodologies for those judgments, e.g., why a model might label something as “fake news,” are difficult to ascertain. We discuss this in Sections 5 and 6.

4 MBFC as described in the literature↩︎

In Section 2, we consider MBFC’s data and methodology but, in Section 3, we show that, despite its methodological flaws, authors still use MBFC widely. We address this tension here in two ways. First, we analyze how authors describe MBFC, arguing that they rarely examine MBFC closely. Second, we ask what authors seek to accomplish with MBFC’s data: What is it that MBFC allows authors to do that makes it so appealing?

4.1 Methods↩︎

We use a combination of examples from our reading and Natural Language Processing (NLP) methods. For the latter, we compile all papers that have MBFC mentions, then, for each paper, we combine the text from the title, abstract, and body of the paper. This text serves as the searchable document for MBFC mentions. Across sentences in these papers, we identify aliases of MBFC (e.g., “MBFC” or “mediabiasfactcheck.com”). When we find these aliases, we build a context window around each sentence containing an MBFC alias that includes 1 sentence before and 1 sentence after.

Table 3: Part of speech top word counts for the context windows mentioning MBFC.
Part of speech Top words
Adjective political, fake, factual, right, low, high, unreliable, extreme, top, reliable, social, left, questionable, independent, different.
Proper noun twitter, facebook, covid, __MBFC_ALIAS__, march, wikipedia, april, alexa.
Noun news, bias, media, source/sources, websites, outlets, domains, information, data, labels, articles, dataset, list, credibility.
Verb used/using/use, left, based, shared, mixed, provided, labeled, obtained, listed, collected, published, classified, rated.

We replace all aliases with the protected token __MBFC_ALIAS__. We de-duplicate contexts, replace digits with NUM, and remove stopwords, including web and academic artifacts (e.g., “http”, “et”, “al”, “fig”, and roman numerals) to produce a bag of words count (see Table 4). Finally, we remove contexts containing a digit directly before an alias, as this most often indicates a footnote, which yield false contexts, since a footnote need not have any relationship to the footnote before or after.

These cleaning steps remove \(241\) contexts, leaving \(728\) for analysis. Note that a single article may have multiple contexts. We tag parts of speech using Stanza [36], then show the most frequent adjectives, proper nouns, nouns, and verbs occurring within our contexts in Table 3.

Table 4: Top words (count) in MBFC-containing articles within the context windows mentioning MBFC. The context windows include a single sentence prior to and after any MBFC alias.
1-19 20-38 39-57 58-75
news (801) articles (116) extreme (76) facebook (57)
__MBFC_ALIAS__ (477) dataset (116) using (75) ratings (57)
bias (351) used (116) number (75) shared (56)
media (337) based (115) sites (74) results (54)
sources (306) high (114) domain (73) social (54)
political (247) credibility (112) twitter (72) far (53)
right (198) website (110) one (71) urls (53)
left (198) tweets (103) mixed (70) scores (53)
fake (181) center (101) top (69) posts (52)
websites (157) leaning (97) use (68) outlet (51)
factual (152) users (96) links (66) biased (51)
source (149) also (95) label (63) questionable (50)
outlets (148) misinformation (95) reliable (62) independent (49)
domains (146) two (90) categories (61) provided (48)
information (145) content (89) analysis (60) online (48)
data (144) pages (89) scale (60) research (48)
labels (134) reporting (83) lists (59) accounts (47)
low (122) score (77) category (58) labeled (47)
list (119) unreliable (76) fact (57)

4.2 Results↩︎

Consistent with our previous work [37], the proper nouns show that studies that mention MBFC focus on social media. In the adjectives, we see many of the descriptors of MBFC’s rating systems discussed in Section 2 ( \(\boldsymbol{\langle}\) low \(\boldsymbol{\rangle}\), \(\boldsymbol{\langle}\) high \(\boldsymbol{\rangle}\), \(\boldsymbol{\langle}\) left \(\boldsymbol{\rangle}\), and \(\boldsymbol{\langle}\) right \(\boldsymbol{\rangle}\) are {sec:credibility}/ {factual reporting} and {bias} scores), and “social” describes its noun-pair “media”. The rest of the nouns contain, similar to the proper nouns, various areas which authors use MBFC to study, like “outlets” and “sources,” as well as MBFC terms like {bias}. In the verbs, there are examples of what MBFC provides the authors, e.g., sources are “rated” or “classified.” Combining these results, we see that studies turn to MBFC to study social media. They use it as a dataset to classify different kinds of sources by their credibility or bias.

From our reading and qualitative sampling of contexts, many texts describe MBFC as a “fact-checking organization” [1][3], [38], often adding the word “independent” [5], [6], or invoking it alongside scholarly sources, e.g., “various scholars and fact-checking organizations” [1]. In total, of the 372 papers that use MBFC, 42 contain the phrase “fact-checking organization”, often describing MBFC directly. Sometimes, texts provide further details, describing MBFC as, for example, “a small independent team of researchers and journalists” [38]. Others provide almost no details at all, and instead solely refer to it only by its URL [7], without using the full name “Media Bias/Fact Check”.

 [5], the most cited paper in our database, gives a typical example that is consistent with our NLP results and our reading:

We tag links as reliable or questionable according to the data reported by the independent fact-checking organization Media Bias/Fact Check. In order to clarify the limits of an approach that is based on labeling news outlets rather than single articles, as for instance performed in [other studies], we report the definitions used in this paper for questionable and reliable information sources. In accordance with the criteria established by MBFC, by questionable information source we mean a news outlet systematically showing one or more of the following characteristics: extreme bias, consistent promotion of propaganda/conspiracies, poor or no sourcing to credible information, information not supported by evidence or unverifiable, a complete lack of transparency and/or fake news. By reliable information sources we mean news outlets that do not show any of the aforementioned characteristics.

This paragraph focuses mostly on what MBFC allows its authors to accomplish, methodologically speaking, but spends very little time discussing MBFC itself. As noted previously, [5] refer to MBFC as an “independent fact-checking organization,” as do many other texts in our database. We rarely get descriptions of MBFC itself beyond a simple sentence like this. Of the most used words in our context windows, only “dataset,” “website,” and “independent” are words that describe MBFC.

The second most cited paper in our database has the same lead author. It describes MBFC with an identical quote and uses it similarly.

In the third-most cited paper, “A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities,” MBFC appears alongside similar datasets in a section titled “Resources for Understanding News Publishers.”

We introduce several resources that can help obtain the ground truth on the credibility (or political bias) of news publishers. One resource is the Media Bias/Fact Check website, which provides a list of media along with their political slant: left, left-center, least biased, right-center, and right. [39]

This example recommends MBFC as a resource. Consistent with our NLP findings and previous examples, we again find no discussion of its methodology, only a description of its data.

As a final example, our fourth most cited paper adds some nuance to this discussion. [7] provide descriptive statistics of their data (which combines MBFC and AllSides, a similar source):

Using this final separation in seven classes, we identify in our dataset (we give the top hostname as an example in parenthesis): 16 hostnames corresponding to fake news websites (e.g. thegatewaypundit.com), 17 hostnames for extremely biased (right) news websites (e.g. breitbart.com), 7 hostnames for extremely biased (left) news websites (e.g. dailynewsbin.com), 18 hostnames for left news websites (e.g. huffingtonpost.com), 19 hostnames for left leaning news websites (e.g. nytimes.com), 13 hostnames for center news websites (e.g. cnn.com), 7 hostnames for right leaning websites (e.g. wsj.com), and 20 hostnames for right websites (e.g. foxnews.com).

They also link to MBFC’s methodology page in their description of MBFC.

The string “mediabiasfactcheck.com/methodology” appears 8 times in our database, which includes the references, making this a rare example of an author acknowledging that MBFC’s data is the result of extensive interpretative work, the nature of which often goes unacknowledged in the literature.

5 Discussion↩︎

5.1 Misconceptions in the literature↩︎

As we discussed in Section 3, though studies often use MBFC’s data, they rarely describe MBFC. When there is a description, it is often inconsistent with how MBFC describes itself. Contradicting many of the examples in Section 4, MBFC’s “About” page contains the following description:

Media Bias Fact Check, LLC is a North Carolina-based Limited Liability Company solely owned and operated by Dave Van Zandt. He makes all final editorial and publishing decisions. [17]

Similarly contradicting previous descriptions, MBFC’s author explicitly denies being a journalist:

Dave Van Zandt is a registered Non-Affiliated voter who values evidence-based reporting. Though not a journalist, Dave has maintained a lifelong interest in politics and media bias. He originally pursued a Communications degree in college before ultimately earning a degree in Physiology. Since then, he has worked in the healthcare industry (Occupational Rehabilitation) while continuing to study media, language, and bias independently.

Over the past 20 years, he has studied media bias and linguistics and has applied the scientific method to create a structured, evidence-based methodology for assessing media bias and factual reporting. [17]

Despite being used as, for example, ground truth in machine learning models [40], or to definitively label sources as unreliable [5], MBFC provides the following disclaimer, which we interpret as being in tension with the previous quote:

Disclaimer: The methodology used by Media Bias Fact Check is our own. It is not a tested scientific method. It is meant as a simple guide for people to get an idea of a source’s bias. [17]

In our database, we have not been able to find a single paper that adequately describes MBFC as the opinions of one person or critically engages with its methodology in order to justify proceeding with its use despite its scientific limitations.

5.2 A critique of MBFC’s methodology↩︎

As we show in the previous section (Section 5.1), there are widespread misconceptions about MBFC in the misinformation literature. Similarly, as we discuss in Section 1, university libraries recommend MBFC, and it is increasingly integrated into consumer products. We therefore consider its methodology here.

In Section 2, we provide descriptions of MBFC’s data, and explain the rubrics MBFC uses in generating ratings. MBFC’s rubrics make clear that its data is the result of extensive interpretive work. This interpretive work, however, does not meet basic academic standards. Our argument here is not that media outlets cannot be quantitatively compared or that measurement is a lossy process. The problem here goes beyond this inherent property of quantification. Put simply, MBFC’s methodology is sloppy.

Recall the arbitrariness in point values found in Section 2. Each score is composed of sub-scores, which are given percentage weights without explanation, e.g., {bias} is 35% {economic system}. Similarly, scores are generated, then transformed into bins of arbitrary sizes, again with no explanation, and only those secondary, transformed values are made available to the end user. The {sec:credibility} rubrics contain nested logic without a theoretical justification. The {bias} values outside the left-right spectrum similarly typify a poorly conceptualized rubric at all levels, from the measurement values themselves to the definitions it uses in its rubrics.

Recall from Section 2 and Table ¿tbl:tab:factual-reporting? that the {factual reporting} score contained within it {one-sidedness}, the best score for which was a lack of bias. The rubric uses the term “bias,” but this use has no explicit relationship to {bias}, which is a seemingly independent measure, though there is clear conceptual slippage between the pair. This means that MBFC’s interpretation of {factual reporting} contains within it a problematic normative claim that centrist reporting is, by definition, more factual. This conceptual slippage goes both ways, as the {bias} score contains values that live entirely outside the {bias} spectrum, but instead seem more like questions of {factual reporting}. In short, though the rubrics that MBFC uses often clearly spell out procedures, these procedures are arbitrary.

In the rare cases that we do get theoretical discussions, they are inadequate. Consider again Table [tbl:tab:economic-system], which contains the rubric for the {economic system} portion of the {bias} score. MBFC’s methodology page gives the appearance of an academic article: It has descriptions, rubrics, mentions of constraints, and a references section at the bottom. Though it satisfies these aesthetic expectations of rigorous work, it deploys these aesthetics without substance. The page contains 18 references at the bottom without in-line citations to indicate from which source(s) any particular definitions are drawn [25]. Upon searching through them, we do not find the MBFC definitions in any of the citations, nor anything topically relevant. These references are styled similar to an academic paper’s references, but they are in fact more like a “further reading” section, and contain no citations supporting the methodology, as one would expect. This lack of citations creates problems of integrity and interpretation for scholarship downstream of MBFC, though it would not be apparent upon skimming the page.

Despite having no meaningful citations, many of the definitions in the {economic system} rubric are contentious, and some have well-documented historical and current counterexamples.

Recall from Table [tbl:tab:economic-system] that the furthest left {economic system} “[a]dvocates no corporatism, extreme regulation, and full government ownership of industries.” Nazi Germany practiced an economic doctrine known as Gleichschaltung (roughly “total coordination”) with the state, requiring that “endless accountings be submitted regularly to government bureaus” and “companies install Hollerith machines [an IBM punch-card computer] to ensure prompt, up-to-the-minute reports” [41]. Fascist Italy underwent similar processes: “State intervention in the economy blurred the lines between the private and public sector to such a degree that employers were in fact transformed into such agents of the state” [42]. Under MBFC’s rubric, these economic systems would classify as leftist. Since {economic system} is 35% of the total {bias} score, it would be mathematically difficult for the final {bias} score of an outlet that championed Mussolini’s Italy or Hitler’s Germany to be further to the right than the center.

For a more current example, roughly 40% of Saudi Arabia’s GDP comes from oil [43] extracted by a state-owned company [44]. According to the rubric, this too is a leftist system, so Saudi propaganda would, we expect, tilt left. MBFC’s ratings for Arab News, which, according to MBFC, is propaganda for Saudi Arabia, rates the source \(\boldsymbol{\langle}\) right-center \(\boldsymbol{\rangle}\). MBFC provides a blurb describing their reasoning, and {economic system}, despite supposedly making up 35% of the score, is not mentioned anywhere on the page [45].

Similarly, MBFC’s data rates the Huffington Post as far left as the scale allows. It is difficult to imagine that Huffington Post advocates anything like “no corporatism, extreme regulation, and full government ownership of industries,” and MBFC’s write-up again makes no mention of Huffington Post as a mouthpiece for the international proleteriat, instead focusing on the outlets attacks on Donald Trump [46].

These examples are reasons to doubt the consistency of application of MBFC’s rubrics, but recall also from Section 2 that we are quoting the new rubric (from 2025). If the rubrics are objectively applied, results from before and after 2025 should not be comparable. Unfortunately, there is no archived version of the pre-2025 data, so quantitative comparisons are impossible.

There is, however, no satisfactory outcome: If the rubrics are objectively and procedurally applied, then studies from before and after are not comparable. If they are not consistently applied, then that is a problem on its face. We have not found a paper in our database that mentions this 2025 rubric change, and the majority of the papers in our database are from before the rubric update, meaning that they were relying on MBFC data before it had even the stylings of rigor.

Setting aside these inconsistencies, as well as historical and extant counterexamples, MBFC’s definitions also invoke yet contradict theoretical traditions without explanation. Communists and socialists have defined these terms for centuries [47]. Rather than a society in which the government owns everything, as MBFC describes it, communism is, according to communists, a stateless society, one that would arrive at the end of a transition period of worker control over the means of production [48], [49]. Communist parties in power have maintained this distinction, viewing their regimes as building towards communism [50], or in the “primary stage” [51].

The right of the spectrum, meanwhile, seems to contain the ideas of the free market from  [52] or  [53] (though, without citations, we cannot be sure). Unlike the dismissive posture taken towards the aforementioned socialist tradition, free market absolutism is presented on its own terms. The rubric does not engage with arguments that markets are not spontaneous, but political creations [54], [55], often coerced into being with military force, and, in the 20th century, often through US intervention, in opposition to democratically-elected governments [56][59].

Recall also, from Section 1, as part of a rightist attack on the media, CMPA deployed a dubious methodology to attempt to propagate the liberal media claim. MBFC’s methodology contains echoes of the same. In the methodology page’s introduction, MBFC writes:

Special attention is given to detecting bias by omission, one-sided narratives, and the use of unreliable sources [17].

We saw this attention in Section 2.4’s {straight news reporting balance}, which...

[m]easures how well a source reports all sides in its straight news stories, either through story selection or content balance within articles. This covers strictly news reporting and is separate from Editorial/Op-Ed bias.

Though not identical, “content balance” echoes CMPA’s “thematic messaging” and other, similar techniques used by various entities discussed in Section 1. We suspect that this is an example of how MBFC has been influenced by much of the history described in Section 1 but, again, without citations, we cannot know.

The point here is not asymmetry, hypocrisy, or incompleteness, but lack of rigor. MBFC presents rubrics absent theory or meaningful, specific citations to the literature. If scholarship is a cumulative endeavor that grows and branches, then scholarship using MBFC develops and belongs on a disconnected branch. MBFC’s own site confirms this observation, as we saw in the disclaimer quoted in Section 5.1.

In the next section, we theorize why, despite methodological flaws, and even a disclaimer, MBFC can be found in studies of media and (mis)information throughout the literature.

5.3 Why studies use MBFC↩︎

When doing their exploratory analysis of MBFC, researchers find data that, presumably, matches what they expect, e.g., that the media is center-left and mostly reliable (Section 2.4), or that official enemies of the US are less free (Section 2.3). They might take this agreement as proof that MBFC is reliable, when instead it is proof that they are both downstream of the same political processes. MBFC’s success in the literature, we argue, comes not from its accuracy, but from its faithful quantification of hegemony. MBFC makes the dominant ideology legible to computational methods.

[60] find that MBFC has a high correspondence with other, similar data sources, such as NewsGuard and Ad Fontes. This is not a coincidence. Each of these different sources, in some form or another, aim for neutrality by sampling average people. In the case of MBFC, \(N=1\) (though MBFC does have volunteers), but Ad Fontes, for example, asks a panel of one conservative, one liberal, and one centrist to rate sources after attending a training [61]. As we saw in Section 1, in politics, “common sense” ideas are downstream of successful political projects, meaning that methods that attempt to be neutral by gauging the general public find hegemony, not neutrality [62].

This problem runs deeper than the final, downstream concepts. We now know the history of the liberal media, but the media is also, in the philosophical sense, a quintessentially liberal institution [19]. It is sometimes called the “Fourth Estate,” a reference to the Estates-General and the French revolution, perhaps the paradigmatic liberal revolution [63]. The “liberal” in liberal media, however, refers to a uniquely American usage of the word, in which there is a center, and “liberal” means left of center. In Australia, by contrast, the Liberal Party is referred to as a conservative party in headlines without controversy [64].

Similar to the battle to label the media as liberal, the meaning of words like “liberal” are active terrains of struggle, as are all words used in politics, because power is wielded through language [62], [65]. For scholarship seeking simple, neutral data on politics [66], this is not a resolvable problem, because, in politics, the definition of every word will be forever contested. New concepts will get mapped onto old words while old concepts get warped, twisted, and discarded as different coalitions jockey for rhetorical position.

As we saw in Table 2, a quarter of the most cited papers leveraging MBFC are attempting to do misinformation detection, or to evaluate the veracity of statements at scale on social media through natural language processing of text. If the stated goal of misinformation research is to intervene in some kind of political malady, as many papers state explicitly [37], then they face a serious challenge, because political struggle happens by and through words. In the process, words and meanings change, and they risk being trapped and confused by these rhetorical maneuvers.

MBFC, however, allows researchers to study politics, if superficially, while actually side stepping it. At first glance, MBFC’s data mirrors a common practice in machine learning, that of human or expert annotation of “ground truth” to train models. Commonly used datasets like MNSIT, for example, provide photographs of hand-written integers [67], and a typical workflow might include human annotation for each image, then training a model to classify new, unannotated photos of integers.

In our case, instead of dealing with inexact language, researchers have a dataset that has, through pseudo-scientific annotation, extracted politics from its native medium. The resulting data would allow them to view politics clearly, using numbers, as if from the outside, rather than bogged down in rhetorical mud. The error is that the inexactness of language is where much political work happens. Outside political struggles, where the definitions of words are less contested, such annotations are on firm ground. When using MBFC, however, in an attempt to climb out of politics, researchers entrench themselves further, inadvertently accepting hegemony. We explore the consequences of this error in our next and final section, where we implore researchers to stop using MBFC.

6 Conclusion↩︎

Figure 5: An illustration of the single point of reliance on MBFC for ground truth on the credibility and bias of news outlets—a ground truth used in various research contexts and commercial products.

As we saw in Section 5.2, MBFC takes for granted the results of a decades-long conservative attack against the media, a propaganda campaign to which the entire Anglosphere has been subjected. MBFC then quantifies this rhetorical victory as data and, aided by the work of misinformation scholars, it is inserted back into discourse as scientific fact.

In a sense, it is true that the media is liberal, but only because it was made true through a conscious political project to discredit it. Now, throughout the misinformation literature, in its attempt to understand and alleviate: “[t]he explosive growth in fake news and its erosion to democracy, justice, and public trust” [68]; “distrust in scientific expertise” [69]; “political partisanship and mistrust of science” [70]; and “distrust in science [that] undermines public health and may drive civil unrest” [71], misinformation scholarship has ironically turned this successful attack into a scientific consensus (each of these research motivation quotes comes from a paper using MBFC). This consensus then buttresses the putative raison d’être of conservative media like Fox News, whose “Fair and Balanced” slogan refers not to its own objectivity, but to its conservatism balancing the liberal media [18].

As we also saw, many of the papers in our database seek to detect misinformation on social media to alleviate a crisis of credibility or institutional trust. In the process of doing so, they unwittingly encode the arguments through which the rightist media ecosystem justifies itself into technical systems that moderate public discourse. In this, misinformation scholars make a mistake with historical precedence. When Ernst wrote his structural press critiques,

...the common sense among postwar media reformers was that the concentration of media ownership would disproportionately stifle liberal voices [...] [W]hen Ernst agreed to join the AIM board [...] the news media now was increasingly considered to be controlled by a cabal of liberals [18].

Ernst’s sincere if naïve commitment to truth in media allowed him to be used by political actors to legitimize themselves. Many journalists made the same mistake:

AIM’s thin veneer of impartiality—achieved in part through the group’s affiliation with Morris Ernst—provided enough plausible deniability that some corner of the journalism profession felt compelled to take it seriously [18].

Misinformation scholarship’s search for a fair, neutral procedure by which to moderate public discourse leaves it vulnerable to cooption by unscrupulous and well-funded political actors, some of whose propaganda campaign it already renders as scientific fact. Just as there is no definition of “liberal” that does not take a political stance, there can be no neutral dataset or detection scheme. So long as truth has partisan opponents, to be for truth is a partisan stance. Ernst’s error, then, should be a lesson in the dangers of naïveté in the face of politics. It cannot be the role of scholarship to feed the output of a political process back into itself, detached from context and laundered through academic prestige. We therefore urge researchers to stop using MBFC.

7 Acknowledgments↩︎

Earlier drafts of this paper predated the publication of AJ Bauer’s wonderful Making the Liberal Media, and, without its aid, the historical discussions of those drafts left much to be desired. We therefore express our gratitude for his work, and congratulate him on the completion of an excellent, thorough, and necessary book.

The authors acknowledge support by National Science Foundation awards #2419830 and #2242829 and MassMutual.

References↩︎

[1]
L. Singh, L. Bode, C. Budak, K. Kawintiranon, C. Padden, and E. K. Vraga, “Understanding high‐ and low‐quality URL sharing on COVID-19 Twitter streams,” Journal of Computational Social Science, vol. 3, pp. 343–366, 2020, doi: 10.1007/s42001-020-00093-6.
[2]
Y. Zhang, L. Wang, J. J. H. Zhu, and X. Wang, “Conspiracy vs science: A large‑scale analysis of online discussion cascades,” World Wide Web, vol. 24, no. 2, pp. 585–606, 2021, doi: 10.1007/s11280‑021‑00862‑x.
[3]
G. Etta, M. Cinelli, A. Galeazzi, C. M. Valensise, W. Quattrociocchi, and M. Conti, “Comparing the impact of social media regulations on news consumption,” IEEE Transactions on Computational Social Systems, vol. 10, 2022.
[4]
M. Cinelli, G. De Francisci Morales, A. Galeazzi, W. Quattrociocchi, and M. Starnini, “The echo chamber effect on social media,” Proceedings of the National Academy of Sciences, vol. 118, no. 9, p. e2023301118, 2021, doi: 10.1073/pnas.2023301118.
[5]
M. Cinelli et al., “The COVID-19 social media infodemic,” Scientific Reports, vol. 10, p. 16598, 2020, doi: 10.1038/s41598-020-73510-5.
[6]
S. Horawalavithana et al., “Vaccination trials on hold: Malicious and low credibility content on Twitter during the AstraZeneca COVID-19 vaccine development,” Comput Math Organ Theory, vol. 29, no. 3, pp. 448–469, Sep. 2023, doi: 10.1007/s10588-022-09370-3.
[7]
A. Bovet and H. A. Makse, “Influence of fake news in Twitter during the 2016 US presidential election,” Nature Communications, vol. 10, no. 1, p. 7, 2019, doi: 10.1038/s41467-018-07761-2.
[8]
G. Wunsch, C. Gourbin, and F. Russo, “Big Data, Demography, and Causality,” Open Journal of Social Sciences, vol. 12, no. 1, pp. 181–206, Jan. 2024, doi: 10.4236/jss.2024.121012.
[9]
K. Odhner, Accessed: 2026-04-03“MediaBiasFactCheck.com as a tool for lateral reading.” Penn State News Literacy Initiative, 2024, [Online]. Available: https://newsliteracy.psu.edu/news/mediabiasfactcheck-com-as-a-tool-for-lateral-reading.
[10]
Wichita State University Libraries, LibGuide; Accessed: 2026-04-03“Evaluating news sources.” Wichita State University Libraries, n.d., [Online]. Available: https://libraries.wichita.edu/c.php?g=613382&p=4263041.
[11]
USC Upstate Library, LibGuide; Accessed: 2026-04-03“Know your news.” University of South Carolina Upstate Library, 2025, [Online]. Available: https://uscupstate.libguides.com/news_aware.
[12]
R. Gallotti, F. Valle, N. Castaldo, P. Sacco, and M. De Domenico, “Assessing the risks of ‘infodemics’ in response to COVID-19 epidemics,” Nature Human Behaviour, vol. 4, no. 12, pp. 1285–1293, Dec. 2020, doi: 10.1038/s41562-020-00994-6.
[13]
M. Gruppi, B. D. Horne, and S. Adalı, NELA-GT-2020: A Large Multi-Labelled News Dataset for The Study of Misinformation in News Articles.” arXiv, Feb. 2021, doi: 10.48550/arXiv.2102.04567.
[14]
H. Lin et al., “High level of correspondence across different news domain quality rating sets,” PNAS Nexus, vol. 2, no. 9, pp. 1–8, 2023, doi: 10.1093/pnasnexus/pgad286.
[15]
J. Nørregaard, B. D. Horne, and S. Adalı, NELA-GT-2018: A Large Multi-Labelled News Dataset for the Study of Misinformation in News Articles,” Proceedings of the International AAAI Conference on Web and Social Media, vol. 13, pp. 630–638, Jul. 2019, doi: 10.1609/icwsm.v13i01.3261.
[16]
Ground News, Accessed: 2026-04-03“Ground news rating system.” https://ground.news/rating-system, 2026, [Online]. Available: https://ground.news/rating-system.
[17]
D. M. Van Zandt, Accessed: 2025-11-05“Media bias/fact check.” Website, 2025, [Online]. Available: https://mediabiasfactcheck.com/.
[18]
A. J. Bauer, Making the liberal media: How conservatives built a movement against the press. New York: Columbia University Press, 2026.
[19]
D. Greenberg, The Idea of ‘the Liberal Media’ and Its Roots in the Civil Rights Movement,” The Sixties, vol. 1, no. 2, pp. 167–186, Dec. 2008, doi: 10.1080/17541320802457111.
[20]
E. S. Herman, The myth of the liberal media: An edward herman reader. New York: Peter Lang, 1999.
[21]
C. J. Stewart, “The master conspiracy of the john birch society: From communism to the new world order,” Western Journal of Communication, vol. 66, no. 4, pp. 423–447, 2002, doi: 10.1080/10570310209374748.
[22]
E. S. Herman and N. Chomsky, Manufacturing consent: The political economy of the mass media. New York: Pantheon Books, 1988.
[23]
P. Hart, “Meet the Myth-Makers.” FAIR, Jul. 01, 1998, Accessed: Jan. 08, 2026. [Online]. Available: https://fair.org/extra/meet-the-myth-makers/.
[24]
E. Grieco, “Cable news fact sheet,” Sep. 14, 2023. https://www.pewresearch.org/journalism/fact-sheet/cable-news/ (accessed Jun. 26, 2026).
[25]
D. V. Zandt, “Methodology,” Media Bias/Fact Check. Jan. 2026, Accessed: Jan. 17, 2026. [Online].
[26]
D. M. Van Zandt, Media Bias/Fact Check: Methodology.” https://web.archive.org/web/20190203144044/https://mediabiasfactcheck.com/methodology/, 2019.
[27]
D. M. Van Zandt, Media Bias/Fact Check: Methodology.” https://web.archive.org/web/20240131160330/https://mediabiasfactcheck.com/methodology/, 2024.
[28]
D. M. Van Zandt, Media Bias/Fact Check: Methodology.” https://web.archive.org/web/20251112074011/https://mediabiasfactcheck.com/methodology/, 2025.
[29]
D. V. Zandt, Accessed: 2026-01-05“Country methodology.” Media Bias/Fact Check; https://mediabiasfactcheck.com/country-methodology/, 2023, [Online]. Available: https://mediabiasfactcheck.com/country-methodology/.
[30]
Enemies of the United States,” World Population Review. Jan. 2026, Accessed: Jan. 26, 2026. [Online]. Available: https://worldpopulationreview.com/country-rankings/us-enemy-countries.
[31]
W. Ammar et al., “Construction of the literature graph in semantic scholar.” 2018, [Online]. Available: https://arxiv.org/abs/1805.02262.
[32]
J. Priem, H. Piwowar, and R. Orr, “OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts,” arXiv preprint arXiv:2205.01833. 2022.
[33]
S. Peroni and D. Shotton, “OpenCitations, an infrastructure organization for open scholarship.” 2019, [Online]. Available: https://arxiv.org/abs/1906.11964.
[34]
Elsevier, Accessed: 2026-04-03“Scopus.” https://www.scopus.com, 2004.
[35]
Google, “Google scholar.” 2026, [Online]. Available: https://scholar.google.com.
[36]
P. Qi, Y. Zhang, Y. Zhang, J. Bolton, and C. D. Manning, “Stanza: A python natural language processing toolkit for many human languages,” in Proceedings of the 58th annual meeting of the association for computational linguistics: System demonstrations, 2020, pp. 101–108, doi: 10.18653/v1/2020.acl-demos.14.
[37]
A. J. Ruiz Iglesias, D. Bennett, J. W. Zimmerman, C. M. Danforth, and P. S. Dodds, “False memories to fake news: The evolution of the term "misinformation" in academic literature.” 2026, [Online]. Available: https://arxiv.org/abs/2602.22395.
[38]
J. Flamino et al., “Political polarization of news media and influencers on Twitter in the 2016 and 2020 US presidential elections,” Nature Human Behaviour, vol. 7, no. 6, pp. 904–916, Jun. 2023, doi: 10.1038/s41562-023-01550-8.
[39]
X. Zhou and R. Zafarani, v2, last revised 17 Jul 2020“A survey of fake news: Fundamental theories, detection methods, and opportunities,” arXiv preprint, vol. arXiv:1812.00315, 2018, [Online]. Available: https://arxiv.org/abs/1812.00315.
[40]
S. Raza and C. Ding, “Fake news detection based on news content and social contexts: A transformer-based approach,” International Journal of Data Science and Analytics, vol. 13, no. 4, pp. 335–362, May 2022, doi: 10.1007/s41060-021-00302-z.
[41]
E. Black, IBM and the holocaust: The strategic alliance between nazi germany and america’s most powerful corporation, expanded edition. Washington, DC: Dialog Press, 2002.
[42]
Z. Elward, Submitted in partial fulfillment of the requirements for the degree of Masters of Arts“Fascist corporativism and the myth of the new state: The construction of the totalitarian state in italy,” Master’s thesis, Central European University, 2023.
[43]
Forbes, Accessed: 2026-03-01Saudi Arabia.” 2023, [Online]. Available: https://www.forbes.com/places/saudi-arabia/.
[44]
Saudi Arabian Oil Company (Saudi Aramco), Archived version; accessed 2026-03-01“Who we are.” 2018, [Online]. Available: https://web.archive.org/web/20180818084004/http://www.saudiaramco.com/en/home/about/who-we-are.html.
[45]
D. M. Van Zandt, Accessed: 2026-06-12“Media bias/fact check: Arab news.” Website, 2025, [Online]. Available: https://mediabiasfactcheck.com/arab-news/.
[46]
D. M. Van Zandt, “Media bias/fact check: Hufftingon post,” 2026. https://mediabiasfactcheck.com/huffington-post/ (accessed Jun. 29, 2026).
[47]
F. Engels, “The principles of communism.” 1847.
[48]
K. Marx, Written in 1875; first published in 1891“Critique of the gotha programme.” 1875.
[49]
[50]
N. P. Trong, “A number of theoretical and practical issues on socialism and the path to socialism in vietnam.” 2021, [Online]. Available: https://vietnamlawmagazine.vn/a-number-of-theoretical-and-practical-issues-on-socialism-and-the-path-to-socialism-in-vietnam-37744.html.
[51]
National People’s Congress of the People’s Republic of China, Constitution of the People’s Republic of China.” 2023, [Online]. Available: http://www.npc.gov.cn/englishnpc/Constitution/node_2825.htm.
[52]
M. Friedman, Capitalism and freedom. Chicago: University of Chicago Press, 1962.
[53]
F. A. Hayek, A free-market monetary system. London: Institute of Economic Affairs, 1976.
[54]
K. Polanyi, The great transformation: The political and economic origins of our time. Boston: Beacon Press, 1944.
[55]
E. P. Thompson, The making of the english working class. London: Victor Gollancz Ltd., 1963.
[56]
V. Bevins, The jakarta method: Washington’s anticommunist crusade and the mass murder program that shaped our world. New York: PublicAffairs, 2020.
[57]
N. Klein, The shock doctrine: The rise of disaster capitalism. Metropolitan Books, 2007.
[58]
D. Graeber, Debt: The first 5,000 years. Brooklyn, NY: Melville House, 2011.
[59]
V. Prashad, Washington bullets: A history of the CIA, coups, and assassinations. New York: NYU Press, 2020.
[60]
H. Lin et al., “High level of correspondence across different news domain quality rating sets,” PNAS Nexus, vol. 2, no. 9, p. pgad286, Sep. 2023, doi: 10.1093/pnasnexus/pgad286.
[61]
Ad Fontes Media, Accessed: 2026-06-09“Methodology.” 2026, [Online]. Available: https://adfontesmedia.com/methodology/.
[62]
A. Gramsci, Selections from the prison notebooks. New York: International Publishers, 1971.
[63]
A. de Tocqueville, L’ancien régime et la révolution. Paris: Michel Lévy Frères, 1856.
[64]
C. Chen, Accessed 2026-01-16Australia’s conservative Liberal Party abandons net zero policy,” Reuters, 2025, [Online]. Available: https://www.reuters.com/sustainability/cop/australias-conservative-liberal-party-abandons-net-zero-policy-2025-11-13/.
[65]
D. Hume, “Of the first principles of government,” in Essays, moral, political, and literary, 1777.
[66]
T. M. Porter, “Thin description: Surface and depth in science and science studies,” Osiris, vol. 27, no. 1, pp. 209–226, 2012, doi: 10.1086/667828.
[67]
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
[68]
M. A. Alonso, D. Vilares, C. Gómez-Rodríguez, and J. Vilares, “Sentiment analysis for fake news detection,” Electronics, vol. 10, no. 11, p. 1348, 2021, doi: 10.3390/electronics10111348.
[69]
J. Lenti et al., Global Misinformation Spillovers in the Vaccination Debate Before and During the COVID-19 Pandemic: Multilingual Twitter Study,” JMIR Infodemiology, vol. 3, p. e44714, May 2023, doi: 10.2196/44714.
[70]
M. Hu, A. Rao, M. Kejriwal, and K. Lerman, Socioeconomic Correlates of Anti-Science Attitudes in the US,” Future Internet, vol. 13, no. 6, p. 160, 2021, doi: 10.3390/fi13060160.
[71]
D. A. Broniatowski, J. R. Simons, J. Gu, A. M. Jamison, and L. C. Abroms, The efficacy of Facebook’s vaccine misinformation policies and architecture during the COVID-19 pandemic,” Science Advances, vol. 9, no. 37, p. eadh2132, 2023, doi: 10.1126/sciadv.adh2132.