And, or course, yet another AI arms race helping to cook the planet.
Some sellers are using AI to fight back against AI. A Chinese toy seller demonstrated to WIRED how they feed refund requests to an AI chatbot to analyze if the photos are doctored.
Well, my continuing work on Bullshit is wearing; Bullshit Fatigue has definitely set in, but I don't think Bullshit Contagion just yet. I try to finish what I start, so, forward we go.
I ran across this book at a good time I think. June 4, 2025, Doc Searls had a great post which I linked to. He linked to the website & online course of our 2 authors:
LOL, the website is https://thebullshitmachines.com/. Are they answering their question? I don't see it right away, but I presume the website pointed at the book.
The book is well-written & easy to read. Maybe a little bit discouraging, because, just as I'm concluding that data & science are possibly our bulwark against the Bullshit Apocalypse, the authors discuss in depth the role of bullshit in data & science. But I think they do finish with hope for the future.
I guess I should take their online class. The book came out in 2020, LLMs didn't start hitting the big time until 2023. Wow, the last time I took an online class was 2014, "The Age of Sustainable Development", Jeffrey Sachs.
Here's the beginning of the book. I think in general, including a book's opening & closing is an appropriate thing to do.
The world is awash with bullshit, and we’re drowning in it.
Politicians are unconstrained by facts. Science is conducted by press release. Silicon Valley startups elevate bullshit to high art. Colleges and universities reward bullshit over analytic thought. The majority of administrative activity seems to be little more than a sophisticated exercise in the combinatorial reassembly of bullshit. Advertisers wink conspiratorially and invite us to join them in seeing through all the bullshit. We wink back—but in doing so drop our guard and fall for the second-order bullshit they are shoveling at us. Bullshit pollutes our world by misleading people about specific issues, and it undermines our ability to trust information in general. However modest, this book is our attempt to fight back.
Hell yeah! Fight back! We need that.
They next:
Talk about Harry Frankfurt's "On Bullshit". Here's my review/summary, I found it overall to be disappointing.
Discuss "bullshit" as a noun vs. a verb. In this review/summary, you have actually 3 glosses for the word "humbug": noun, verb, noun. For "bullshit", the 3 glosses would be "bullshit", "bullshit", "bullshitter". I suspect there is insight to be gleaned from making this distinction, but I'm not at all sure, my subconscious is still on the case.
They define "old-school bullshit" as "rhetoric or fancy language", and give examples, blah, blah, blah.
They define "new-school bullshit":
New-school bullshit uses the language of math and science and statistics to create the impression of rigor and accuracy. Dubious claims are given a veneer of legitimacy by glossing them with numbers, figures, statistics, and data graphics.
The book shares its name with a course the 2 authors give at the U of Washington.
They raise an interesting point: that STEM education doesn't teach bullshit detection & countermeasures like liberal arts does.
Over a century ago, the philosopher John Alexander Smith addressed the entering class at Oxford as follows:
Nothing that you will learn in the course of your studies will be of the slightest possible use to you [thereafter], save only this, that if you work hard and intelligently you should be able to detect when a man is talking rot, and that, in my view, is the main, if not the sole, purpose of education.
For all of its successes, we feel that higher education in STEM disciplines—science, technology, engineering, and mathematics—has dropped the ball in this regard.
...
For a number of reasons, we draw heavily on examples from research in science and medicine in the chapters that follow. We love science and this is where our expertise lies. Science relies on the kinds of quantitative arguments we address in this book. Of all human institutions, science seems as though it ought to be free from bullshit—but it isn’t.
So this book will deal mostly with bullshit in science. Say it ain't so! But still ...
For all our complaints, for all the biases we identify, for all the problems and all the bullshit that creeps in, at the end of the day science works.
This is a book about bullshit. It is a book about how we are inundated with it, about how we can learn to see through it, and about how we can fight back. First things first, though. We would like to understand what bullshit is, where it comes from, and why so much of it is produced. To answer these questions, it is helpful to look back into deep time at the origins of the phenomenon.
Bullshit is not a modern invention.
The Greek Sophist philosophers are identified as some of the founding fathers of bullshit - another word for my list of "other words for bullshit": sophistry.
Some interesting information, talking about mantis shrimp, normally formidable predators, "bluffing" when they are molting. I had identified "bluffing" as a form of bullshit. I think lots of other species bluff - like when they puff themselves up to make themselves appear larger than they are. I think the authors are identifying bluffing as maybe bullshit v0.5.
bluffing does feel rather like a kind of bullshit—but it’s not very sophisticated bullshit.
...
A sophisticated bullshitter needs a theory of mind — she needs to be able to put herself in the place of her mark.
Next up: corvids, FTW! I luv corvids! Corvids may be the only non-human animals known to have a theory of mind.
So when a raven pretends to cache a snack but is actually just faking, does that qualify as bullshitting? ... Full-on bullshit is intended to distract, confuse, or mislead—which means that the bullshitter needs to have a mental model of the effect that his actions have on an observer’s mind.
So, corvids have a theory of mind, but the next species we look at - humans - "take bullshit to the next level".
Human language is such an incredibly powerful tool. It is a generative grammar, which means there is absolutely no limit to how complex a sentence you can form: you can alway add another adjectival phrase, another conjunction. We evolved our big brains to be able to process longer & longer sentences.
It is so incredible that it developed via a sexual selection arms race: the jive wars: males telling females "Oh, baby, you so fine, you're the only one for me!" & females replying, "Oh shut up, Leroy, you so full of shit!". It broke my mind for years when I realized that.
But, having such an incredible communication mechanism is, in the authors' words, "a two-edged sword". It gives all humans a powerful channel into other humans' minds.
So why is there bullshit everywhere? [My ordered list]
Part of the answer is that everyone, crustacean or raven or fellow human being, is trying to sell you something.
Another part is that humans possess the cognitive tools to figure out what kinds of bullshit will be effective.
A third part is that our complex language allows us to produce an infinite variety of bullshit.
The next section is titled "Weasel Words and Lawyer Language". Another word for the list: paltering.
If I deliberately lead you to draw the wrong conclusions by saying things that are technically not untrue, I am paltering.
Bill Clinton's testimony in his impeachment trial is given as the textbook example. Man, if the Union of Concerned Linguists had a doomsday clock counting down to the Bullshit Apocalypse, Clinton's testimony would have moved the hands forward by a minimum of several hours. It was truly a landmark in the history of bullshit.
There is a good bit of analysis about this statement made about a coworker John:
“John doesn’t shoot up when he is working.”
So I find myself calling this "damning with faint praise" - no that's not quite it. "Backhanded compliment"? No, not that either. I know there's a phrase that better describes this, but I'm not getting it. But still, 2 or 3 idioms that click in response to a common form of misinformation? It shows how bullshit is entwined within our language & culture.
Within linguistics, this notion of implied meaning falls under the area of pragmatics. Philosopher of language H. P. Grice[1913-1988] coined the term implicature to describe what a sentence is being used to mean, rather than what it means literally.
...
But implicature is also what lets us palter.
...
An important genre of bullshit known as weasel wording uses the gap between literal meaning and implicature to avoid taking responsibility for things. This seems to be an important skill in many professional domains.
LOL! No kidding? Advertising, politics, public relations, lawyers - they bullshit like they breath.
corporate weaselspeak diffuses responsibility behind a smoke screen of euphemism and passive voice.
I wonder if their discussion of "signaling" comes from semiotics? They talk about:
"self-regarding signals", which is most of what all non-humans do, signal about their state.
other-regarding signals, which "refer to elements of the world beyond the signaler itself".
But even when humans are ostensibly communicating about elements of the external world, they may be saying more about themselves than it seems.
So many forms of bullshit: tribal ritual, recreational, telling tall tales, the jive wars - the list goes on and on. All of these have been amplified by the Internet.
But surely truth will win out in the end? Oops, the next section is titled "Falsehood Flies and the Truth Comes Limping After", from Jonathan Swift, 1710. We are introduced to an important principle:
Perhaps the most important principle in bullshit studies is Brandolini’s principle. Coined by Italian software engineer Alberto Brandolini in 2014, it states:
“The amount of energy needed to refute bullshit is an order of magnitude bigger than [that needed] to produce it.”
...
A few years before Brandolini formulated his principle, Italian blogger Uriel Fanelli had already noted that, loosely translated, “an idiot can create more bullshit than you could ever hope to refute.”
We next get some examples:
Wakefield's fraudulent research that, despite complete & utter debunking, is still a mainstay of the antivax movement - which, unbelievably, is now setting public health policy for the US.
The tale of a Sandy Hook school shooting survivor being killed by the Boston Marathon bomb - complete bunkum.
Truth telling is losing the race against bullshit, & badly. And technology is hurting the cause, not helping it.
This chapter starts out talking about what I have many times declared to be the wonder of our age: the smartphone. It's a communicator! It can access most of the knowledge in the world! It's a calculator! A flashlight! A guitar tuner! Surely, the access to all this information will put an end to bullshit. Nope.
smartphones have become just one more vehicle for spreading bullshit.
...
Technology didn’t eliminate our bullshit problem, it made the problem worse.
New technology always has opponents decrying it. It happened with the printing press; with television & other modern media.
The latest iteration in communications technology, the Internet, is a completely different beast. It has democratized information distribution. This is good - I can write & share a blog, music videos, whatever. But it also is bad - I can be writing complete & utter bullshit, I can be acting as any of a number of kinds of malicious agents, maybe even on behalf of a corporate agenda or a hostile foreign power.
Back when news arrived at a trickle, we might have been able to triage this information effectively. But today we are confronted with a deluge.
...
Because there is so much more volume and so much less filtering, we find ourselves like the Sorcerer’s Apprentice: overwhelmed, exhausted, and losing the will to fight a torrent that only flows faster with every passing hour.
LOL, they just identified Bullshit Fatigue!
Plus, if it was a "deluge" & we were "overwhelmed" before generative AI, where are we at now, with AI slop rapidly outpacing human slop?
So much of the Internet is funded by the attention economy, you have to get clicks above all else.
Quality of information and accuracy are no longer as important as sparkle. A link needs to catch your eye and pull you in. Internet publishers are not looking for Woodward and Bernstein. Instead, they want “Seven Cats That Look Like Disney Princesses,” ...
...
Click-driven media ... drives an arms race among headlines.
This is interesting. What makes the best headlines?
The study [100 million articles published in 2017] found that the most successful headlines don’t convey facts, they promise you an emotional experience. The most common phrase among successful Facebook headlines, by nearly twofold, is “will make you,” as in “will break your heart,” “will make you fall in love,” “will make you look twice,” or “will make you gasp in surprise” as above.
They describe how clickbait headlines gradually got adopted even by respectable outlets. Gawd, I hate clickbait headlines! I try very hard to not click them.
Moving to broadcast media, the authors remind us of the 1987 repeal of the FCC's Fairness Doctrine under St. Reagan. 1 year later we had Faux "News", and the race to partisan news broadcasting was on. Faux wasn't even enough, they also mention "hyperpartisan" news like Breitbart.
Publishers churn out partisan and hyperpartisan content because it pays to do so. Social media favors highly partisan content. It is shared more than mainstream news, and once shared, it is more likely to be clicked on. Deepening the ideological divide has become a lucrative business.
Network news used to be "if it bleeds, it leads". Social media is even worse - the more a post inflames emotions, the more it is upvoted by the algorithms.
We are introduced to MIT Professor Judith Donath and communication theory: communication is not just about conveying information detachedly: it is also about signalling membership in a community - note the same root, "the Latin verb communicare, “to make shared or common.”"
The (online) communities naturally develop "tribal epistemologies in which the truth itself has less to do with facts and empirical observation than with who is speaking and the degree to which their message aligns with their community’s worldview."
The purpose of social media algorithms that determine your feed surely is to keep us informed on topics in which we are interested, yes? No.
These algorithms are not designed to keep you informed; they are designed to keep you active on the platform.
It's all about keeping those clicks, & the associated ad $$$, coming.
the algorithms driving social media content are bullshitters. They don't care about the messages they carry. They just want our attention and will tell us whatever works to capture it.
Social media is fertile ground for both:
"misinformation—claims that are false but not deliberately designed to deceive." Wanting to be 1st with a story drives a lot of this.
"disinformation, falsehoods that are spread deliberately."
Wow, I had not heard of this: in December 2016, a fake news site said that Israel was threatening Pakistan with nuclear weapons. None other than the Defense Minister of Pakistan got taken in & issued a counter threat to Israel.
A single fake news piece led one major power to threaten another with a nuclear attack.
But it's not just fake news; it's also targetted propaganda. The targetting is not concise and controlled, it is more like a shotgun - the authors call it the "firehose strategy". It is deliberately trying to induce Bullshit Fatigue.
In 2016, chess grand master Garry Kasparov summarized this approach in a post on Twitter: “The point of modern propaganda isn’t only to misinform or push an agenda. It is to exhaust your critical thinking, to annihilate truth."
But, it always comes back to the same thing: $$$. Capitalism FTL.
Still, fake news is not primarily a propaganda tool. Most fake and hyperpartisan news is created for a different reason: to generate advertising revenue.
We are told about teenagers in Macedonia running a lucrative business built on fake news. [A recent news story talked about how a large number of MAGA influencer accounts on X/Twitter are coming from foreign countries.] Get those clicks, get the $$$.
The final section is titled "The New Counterfeiters".
In an Internet-connected world, governments have to worry about a new kind of counterfeiting—not of money, but of people. Researchers estimate that about half of the traffic on the Internet is due not to humans, but rather “bots,” automated computer programs designed to simulate humans.
We are given some examples, including more than just bots:
A 2017 bot campaign to oppose Net Neutrality to the FCC: millions of phony messages.
An all-American girl who was actually a Russian propaganda outfit.
Deepfakes, both of images and voices.
The chapter concludes with "three basic approaches for protecting ourselves against misinformation and disinformation online."
Technology. "Tech companies might be able to use machine learning to detect online misinformation and disinformation." Yeah right. We know where this is heading:
Technologically, the same artificial intelligence techniques used to detect fake news can be used to get around detectors, leading to an arms race of production and detection that the detectors are unlikely to win.
Governmental regulation. The authors discuss some of was going on back then, I don't know how relevant that is now. They raise 2 problems: 1) the 1st amendment; 2) who gets to decide what is fake news? With the current Orange Turd administration, government regulation of media is totally the road to perdition.
This is a great idea:
We would like to see users control the information that comes across their social media feeds, rather than being forced to rely on a hopelessly opaque algorithm.
But good luck with that. I will stick with blogs: you subscribe, you get the content. No bribable algorithm. They mention something like the FCC Fairness Doctrine for social media - good luck with that.
Education.
If we do a good job of educating people in media literacy and critical thinking, the problem of misinformation and disinformation can be solved from the bottom up. That is our focus in this book, and in much of our professional lives.
I learned a lot from this book, I think most people could.
I have a working definition that I like not much at all.
The authors note that it can be a catch-all kind of word. I think this is related to Bullshit Contagion.
They note that pleasantries & social niceties are "are often bullshit, but they’re not really the kind of bullshit we’re concerned with here."
[Time for a musical break, suggested by the current subject matter!
]
Insincere promises and outright lies get a bit closer to the mark. ... Still, we tend to think of these claims as outright lies rather than bullshit.
But lies are often most persuasive when dressed in superfluous details; these details come quite close to what we mean by “bullshit.”
They offer their take on Harry Frankfurt's definition:
He described bullshit as what people create when they try to impress you or persuade you, without any concern for whether what they are saying is true or false, correct or incorrect.
This next seems like the maybe most extreme form of bullshit:
Bullshit can be total nonsense. Another philosopher to take up the issue of bullshit, G. A. Cohen[1941-2009], notes that a lot of bullshit—particularly of the academic variety—is meaningless and so cloaked in rhetoric and convoluted language that no one can even critique it. Thus for Cohen, bullshit is “unclarifiable unclarity.” Not only is the bullshitter’s prose unclear, but the ideas underlying it are so ill-formed that it cannot possibly be clarified. Cohen suggests a test for unclarity: If you can negate a sentence and its meaning doesn’t change, it’s bullshit.
They mention "persuasive bullshit" and "evasive bullshit" - pretty obvious that these are. They then offer up their definition of bullshit:
Bullshit involves language, statistical figures, data graphics, and other forms of presentation intended to persuade or impress an audience by distracting, overwhelming, or intimidating them with a blatant disregard for truth, logical coherence, or what information is actually being conveyed.
We learn about sociologist of science Bruno Latour [1947–2022]:
Latour looks at the power dynamics between an author and a reader. In Latour’s worldview, a primary objective of nonfiction authors is to appear authoritative.
This is where the book starts to turn more towards bullshit in science rather than bullshit in general.
According to Latour, scientific claims are typically built upon the output of metaphorical “black boxes".
Most forms of experimentation which provide the data that is the lifeblood of science involve specialized equipment, and specialized expertise. It is easy to use these black boxes as a curtain to hide behind. There are some really good examples given here. This is one of this book's best features: clear, easy-to-understand examples.
More on lies vs. bullshit:
This is where lies and bullshit come together: In our view, a lie becomes bullshit when the speaker attempts to conceal it using various rhetorical artifices.
This is I think an important point, I think the lowest level of bullshit in science:
If the data that go into the analysis are flawed, the specific technical details of the analysis don’t matter.
This chapter ends with an in-depth look as a research paper that reached unbelievably wrong conclusions:
“Automated Inference on Criminality Using Face Images”
Analyzing pictures of criminals vs non-criminals, they found signifigant differences. Spoiler alert - the difference was that the criminals normally aren't smiling in their pic! LOL!
The authors give us a worthwhile heuristic from this.
You might find yourself thinking that even without opening the black box, this type of analysis takes considerable time and focus. That’s true. But fortunately, some claims should more readily trigger our bullshit detectors than others. In particular, extraordinary claims require extraordinary evidence.
In this chapter, we will show you how to think rigorously about associations, correlations, and causes—and how to spot bullshit claims that confuse one for the other.
We get introduced to some Statistics 101: linear correlations and correlation coefficients. The latter range from 1 to -1 - perfect correlation to perfect anti-correlation, with values around 0 meaning little or no correlation.
Let's cut to the chase:
It is a truism that correlation does not imply causation.
...
Unfortunately, one of the most frequent misuses of data, particularly in the popular press, is to suggest a cause-and-effect relationship based on correlation alone. This is classic bullshit, in the vein of our earlier definition, because often the reporters and editors responsible for such stories don’t care what you end up believing.
...
At best, they are trying to tell a good story. At worst, they are trying to compel you to buy a magazine or click on a link.
After I read this section, it seemed like immediately every story I read in the local newspapers that involved data did indeed claim causation from correlation. "x leads to y". "a makes you prone to b". The effect did seem to wear off after a while.
The authors give many examples of this.
They go into depth in a well-known example that my wife has recounted many times over the years: the marshmallow test. Children who delay gratification to get a bigger reward are more successful adults. Because self-control is imporant to success? No. The children who went on & ate the marshmallow were generally from poorer families, where a marshmallow in the hand was definitely worth 2 in the bush. And do we doubt that children of affluent families are primed to be more successful adults? LOL!
This is an interesting point that the marshmallow test illustrates:
But if one is not careful, looking at events chronologically can be misleading. Just because A happens before B does not mean that A causes B—even when A and B are associated. This mistake is so common and has been around for so long that it has a Latin name: post hoc ergo propter hoc. Translated, this means something like “after this, therefore because of it.”
In a footnote, the authors give us I think a userful term:
Statisticians sometimes use the term confounding to refer to situations where a common cause influences two variables that you are measuring.
The section "Spurious Correlations" covers something I had not heard of: data dredging.
Basically, if you randomly compare lots of different sequences of numbers, you will of necessity uncover many, many meaningless correlations. The authors again give great & funny examples.
In the section "Smoking Doesn't Kill?", the authors give us some useful definitions of causality:
There is a key distinction between a probabilistic cause (A increases the chance of B in a causal manner), a sufficient cause (if A happens, B always happens), and a necessary cause (unless A happens, B can’t happen).
There is a hilarious analysis of some "bullshit of a higher grade than usually appears in print" by then vice-president Mike Pence. Wow, those were different times. Pence seems to be at least trying to make sense, unlike the outright pandering we get from the Orange Turd administration - a career politian vs Faux "News" hosts.
In the last section of the chapter, the authors mention an important concept, which I have run into before and had identified as a cognitive illusion or error: selection bias. That gets its own chapter in a bit.
Numbers are ideal vehicles for promulgating bullshit. They feel objective, but are easily manipulated to tell whatever story one desires. Words are clearly constructs of human minds, but numbers? Numbers seem to come directly from Nature herself. We know words are subjective. We know they are used to bend and blur the truth. Words suggest intuition, feeling, and expressivity. But not numbers. Numbers suggest precision and imply a scientific approach. Numbers appear to have an existence separate from the humans reporting them.
But, like in the joke that starts this chapter, about a mathematician, an engineer, & an accountant asked how much "2+2" is, and the accountant answers "What do you want it to be?" What's the old saying: there's lies, damn lies, and statistics" - attributed to Mark Twain, who attributed it to Benjamin Disraeli.
The authors now discuss methods of counting and of creating statistics. It is the normal case that one cannot exhaustively count or measure to get a number - so one uses (random) samples.
We devote the rest of the chapter to considering the subtle ways in which a sample can turn out to be uncharacteristic of the population.
In these examples, we are observing a population with a range of values—a range of heights, for instance—and then summarizing that information with a single number that we call a summary statistic.
They give 1 example (man, their examples are so great, so easy to understand) of a tax cut that saves "the average American" $4000/year. But the average American actually gets $0 tax cut - the tax cuts all go to the top 1%, but when you take the average across the whole population, the number comes out $4000. Grrr.
There are many ways for error to creep into facts and figures that seem entirely straightforward. Quantities can be miscounted. Small samples can fail to accurately reflect the properties of the whole population. Procedures used to infer quantities from other information can be faulty. And then, of course, numbers can be total bullshit, fabricated out of whole cloth in an effort to confer credibility on an otherwise flimsy argument. We need to keep all of these things in mind when we look at quantitative claims. They say the data never lie—but we need to remember that the data often mislead.
A section titled "Pernicious Percentages" starts with a reference to someone who was definitely an early bullshit hunter:
The twelfth chapter of Carl Sagan’s 1996 book, The Demon-Haunted World, is called “The Fine Art of Baloney Detection.”
LOL, another word for my list of "words for bullshit": baloney.
They discuss a technical issue that has always bothered me: percentage increases and decreases. If something doubles, it is a 100% increase, but if it halves, it is a 50% decrease??? It's never seemed right to me. They present several examples of this confusion being exploited to mislead.
While Goodhart’s original formulation is a bit opaque, anthropologist Marilyn Strathern rephrased it clearly and concisely:
When a measure becomes a target, it ceases to be a good measure.
LOL, Goodhart's orIginal is indeed opaque. Per Wikipedia:
Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.
I have family members who are elementary school teachers, & all say that having to "teach to the test" significantly cuts into their time to really give the students the understanding they need.
The word of the year in 2006 was "truthiness". [The current word of the year was just announced; it is "slop", as in generative AI slop.]
The term, coined in 2005 by comedian Stephen Colbert, is defined as “the quality of seeming to be true according to one’s intuition, opinion, or perception without regard to logic, factual evidence, or the like.” In its disregard for actual logic and fact, this hews pretty closely to our definition of bullshit.
We propose an analogous expression, mathiness. Mathiness refers to formulas and expressions that may look and feel like math—even as they disregard the logical coherence and formal rigor of actual mathematics.
OMG, the examples they show gave me bad flashbacks to my corporate years. Some marketoid comes up with some complete bullshit formula with quantities being added, subtracted, multiplied, & divided. The authors make a good point: adding or mulsiplying show that the 2 quantities complement each other; subtracting or dividing means they cancel each other. But THE FORM OF THE EQUATION IS USUALLY COMPLETE BULLSHIT. Why not just say, positive or negative influence, instead of going for "mathiness"?
You can claim your description wasn’t meant to be taken literally, but why not say what you mean in the first place?
The chapter ends with a section on "Zombie Statistics".
Zombie statistics are numbers that are cited badly out of context, are sorely outdated, or were entirely made up in the first place—but they are quoted so often that they simply won’t die.
They follow an example through its history to reach the conclusion: it is completely bogus. It is bullshit.
[Hmmm, Paul Krugman, 1 of my heroes, has been talking about Zombie Economics & Zombie Ideas for years. He had a book published in January 2020, close to this book: "Arguing with Zombies: Economics, Politics, and the Fight for a Better Future". I'm guessing there is a Venn diagram with Zombie Statistics & Zombie Ideas overlapping.]
This chapter opens with a funny anecdote: skiing at a resort for the 1st time, a young author thinks it's amazing that everyone he talks to touts this resort over better known, more popular resorts. His dad straightens him out: of course they like this resort, they are here, aren't they?
This is our 1st example of selection bias, which we heard about a couple of chapters back. This example is "duh!", we will see others that aren't so much.
we introduced the notion of statistical tests or data science algorithms as black boxes that can serve to conceal bullshit of various types. We argued that one can usually see this bullshit for what it is without having to delve into the fine details of how the black box itself works. In this chapter, the black boxes we will be considering are statistical analyses, and we will consider some of the common problems that can arise with the data that is fed into these black boxes.
We return to the topic of sampling, of using "random" samples to generate a statistic.
The problem with this approach [sampling] is that what you see depends on where you look.
Interesting, using autocomplete as an oracle returns [or did in 2019 - maybe not as much payola going on then] completely different results on Facebook vs Google: Facebook is much more positive, people go there to brag, Google is where people go with problems, so it is much more negative.
We should stress that a sample does not need to be completely random in order to be useful. It just needs to be random with respect to whatever we are asking about.
I think I had heard this next before - talk about a non-random sample.
One aim of social psychology is to uncover universals of human cognition, yet a vast majority of studies in social psychology are conducted on what Joe Henrich and colleagues have dubbed WEIRD populations: Western, Educated, Industrialized, Rich, and Democratic. Of these studies, most are conducted on the cheapest, most convenient population available: college students who have to serve as study subjects for course credit.
The authors' definition of "selection bias":
Selection bias arises when the individuals that you sample for your study differ systematically from the population of individuals eligible for your study.
They give lots of examples.
“New GEICO customers report average annual savings over $500” on car insurance.
Well, switching car insurance is a pain. So you're only going to do it if you save a fair amount. LOL, all the car insurance companies say the same thing, another giveaway.
A subset of selection bias, with several examples given:
observation selection effects ... are driven by an association between the very presence of the observer and the variable that the observer reports.
I had not heard of Berkson's paradox. LOL, this plot shows "Why are hot guys such jerks?" through the lens of the paradox.
I guess I would summarize the paradox as, don't expect to make correlations between different attributes in a population that has already been heavily preselected, perhaps implicitly, for those attributes. The Wikipedia article suggests that it is based on people having no intuitive feel for conditional probabilities? LOL, remembering my statistics class 55 years ago, I remember those were important.
A variation on selection bias:
data censoring ... occurs when a sample may be initially selected at random, without selection bias, but a nonrandom subset of the sample doesn’t figure into the final analysis.
2 examples are given, both of which involve life expectancy. In both, the samples are right-censored: subjects that outlive the survey period are effectively removed from the sample! That ain't gonna work!
The final section is "Disarming Selection Bias". The example they analyze is employer wellness programs. They showed how healthy individuals tend to self-select to participate, which lead to the conclusions that such programs worked, because the participants were healthier than average! They showed how randomly selecting who was allowed in the programs demonstrated that the programs were worthless.
This chapter is totally chock-full of interesting information!
It starts with an example of a terribly misleading chart: a chart of Gun Deaths in Florida (y axis) vs time (x axis). It shows a huge drop when Florida's Stand Your Ground took effect in 2005. But, the y axis is reversed!
the graphic designer explained her thought process in choosing an inverted vertical axis: “I prefer to show deaths in negative terms (inverted).”
So a bad design decision, which leads us to an important "calling bullshit" principle:
Never assume malice or mendacity when incompetence is a sufficient explanation, and never assume incompetence when a reasonable mistake can explain things.
In the section "The Dawn of Dataviz" we get a history of the use of graphs & charts:
In the late 18th & early 19th century, demographer William Playfair[1759–1823] pioneered the forms of data visualization that Microsoft Excel now churns out by default: bar charts, line graphs, and pie charts.
Early in that period, physical scientist Johann Heinrich Lambert[1728-1777] published sophisticated scientific graphics of the sort we still use today.
Chart use greatly increased with the advent of digital plotting software around 1980.
But, "A recent Pew Research Center study found that only about half of Americans surveyed could correctly interpret a simple scatter plot."
Edward Tufte [1942-present] was a statistician & political scientist "noted for his writings on information design and as a pioneer in the field of data visualization". In 1982, he (self-)published his landmark book "The Visual Display of Quantitative Information". I had a copy of it, I think I gave it my designer oldest daughter Erica around her college years in the late 90's. Tufte & his book will come up a few times in the chapter.
LOL, the section titled "Duck!" gives an surprising history to a term used in data visualization: duck. In Flanders, NY, there is a building, built in 1931 by a duck farmer to sell ducks & eggs, shaped like a duck!
In architecture, the term “duck” refers to any building where ornament overwhelms purpose, though it is particularly common in reference to buildings that look like the products they sell.
...
Edward Tufte pointed out that an analogous problem is common in data visualization. ... Graphs that violate this principle are called “ducks."
Who knew?
"USA Today was among the pioneers of the dataviz duck." Not surprising, a crapulous newspaper has crapulous charts.
But, of course, they had lots of company. Several examples of unbelievably bad charts are given.
Q: What makes ducks so bad?
A: "the attempt to be cute makes it harder for the reader to understand the underlying data."
Man, "duck" is bad enough as a name, the next chart bullshit we learn about has a horribly bloody-minded name. Who came up with these?
in the original Grimm brothers’ version of the tale [Cinderella], the evil stepsisters make desperate attempts to fit into the glass slipper. They slice off their toes and heels in an effort to fit their feet into the tiny and unyielding shoe.
If a data visualization duck shades toward bullshit, a class of visualizations that we call glass slippers is the real deal. Glass slippers take one type of data and shoehorn it into a visual form designed to display another. In doing so, they trade on the authority of good visualizations to appear authoritative themselves. They are to data visualization what mathiness is to mathematical equations.
OMG, the 1st "glass slipper" they give is indeed classic: "the periodic table of {blah}".
Mendeleev's Periodic Table of the Elements was published in 1867.
His efforts were a triumph of data visualization as a tool for organizing patterns and generating predictions in science. The periodic table is an arrangement of the chemical elements from lightest to heaviest. The left-to-right positions reflect what we now understand to be the fundamental atomic structure of each element, and predict the chemical interactions of those elements. The particular blocky structure of the periodic table reflects the way in which electrons fill the electron subshells around the atomic nucleus. By laying out the known elements in a way that captured the patterns among them, Mendeleev was able to predict the existence and properties of chemical elements that had not yet been discovered.
And, the elements he predicted were indeed discovered! Science! FTW!
So the periodic table is in reality good for 1 thing: showing the structurally-based familes of the elements.
LOL, I searched my phone photos for "periodic table" and only had 1: the real periodic table. I know I had a periodic table of food at some point (bacon was element 1). And a periodic table of Marvel superheroes. And lots of others. The authors totally have me beat:
We’ve seen periodic tables of cloud computing, cybersecurity, typefaces, cryptocurrencies, data science, tech investing, Adobe Illustrator shortcuts, bibliometrics, and more. Some, such as the periodic table of swearing, the periodic table of elephants, and the periodic table of hot dogs, are almost certainly tongue in cheek.
And many, many more, finally terminating with:
Fortunately, someone has created a periodic table of periodic tables.
Next up, subway maps. Great for showing subways. Probably not so great for many other types of data shoehorned glass-slippered into the format.
Periodic tables and subway maps are highly specific forms of visualization. But even very general visualization methods can be glass slippers. Venn diagrams, the overlapping ovals used to represent group membership for items that may belong to multiple groups, are popular glass slippers.
Several examples of laughably bad Venn diagrams follow. They go into one of the common forms popular with marketoids, 3 overlapping circles, with the 3 x 2 circle overlapping regions & the 1 x 3 circle overlapping region being labeled with meaningless bullshit. Man, again, reading this crap takes me back to being corporate.
Another popular form of diagram, particularly in fields such as engineering and anatomy, is the labeled schematic.
I'm going to include their 1st example of labeled schematic bullshit. Again, this has "marketoid" written all over it.
The next section is titled "An Axis of Evil". This is a particular favorite of Faux "News". I bet they have dedicated personnel who are masters of creating misleading axes. I think in this section, we go from dataviz that is confusing or stupid to dataviz that is intentionally dishonest.
Many data graphics, including bar charts and scatter plots, display information along axes. These are the horizontal and vertical scales framing the plot of numeric values. Always look at the axes when you see a data graphic that includes them.
They 1st discuss what I would think is the most commonly used form of misdirection: axes, particulary the y-axis, that don't go to 0. A change of a few % in a # can be turned into 100s of % with proper y-axis manipulation.
They give another example of a bar graph that is just plain cheating: the bar heights don't match the numbers printed in them.
Another example they show is a horrible chart of Annual Global Temperatures from 1880 to 2020. The y-axis goes from -10 degrees to 110 degrees. Why, there's no change at all!
But, replot the data with the y-axis going from 55 to 60 degrees. The 1 degree change is readily visible.
The disingenuous aspect of the ... graph is that Hayward made graphical display choices that are inconsistent with the story he is telling.
This next again is talking about straight-up mendacity:
We have to be even more careful when a graph uses two different vertical axis scales. By selectively changing the scale of the axes relative to each other, designers can make the data tell almost any story they want.
The examples here are so instructive. This kind of data manipulation makes me angry. I was trained as a scientist (physics), and this is the opposite of science.
Another heuristic, from an amazing example where they have a y-axis for "Glyphosate applied (1,000 tons)" that goes to -10! How do you apply -10,000 tons of something!
We’ve noted that the vertical axis need not go to zero for a line graph, but if it goes to a negative value for a quantity that can take on only positive values, this should set off alarm bells.
More principles, things to watch for:
Make sure that the time frame depicted is appropriate for the point the graph is meant to illustrate.
...
In general, we need to be on the lookout for uneven or varying scales on the x axis. Something similar can happen with bar charts, when data are “binned” together to form bars.
The final section of this chapter is titled "The Principle of Proportional Ink". Here's the principle:
When a shaded region is used to represent a numerical value, the size (i.e., area) of that shaded region should be directly proportional to the corresponding value.
This is a specialization of a principle from Tufte's book:
the representation of numbers, as physically measured on the surface of the graphic itself, should be directly proportional to the numerical quantities represented.
This leads us to some maybe slightly unintuitive, heavy but powerful stuff:
we explained how a bar graph emphasizes magnitudes, whereas a line graph emphasizes the changes. As a result, a bar graph should always have a baseline at zero, whereas a line graph is better cropped tightly to best illustrate changing values. Why the apparent double standard?
The principle of proportional ink provides the answer. This principle is violated by a bar chart with axes that fail to reach zero.
The examples that follow clearly illustrate this principle.
Here's another derivative principle:
a “filled” line chart, which does use shaded areas to represent values, should have an axis that goes to zero.
The "donut bar chart" is identified as violating the principle.
This next I was surprised by, but it makes total sense: most 3d graphics are bogus & violate the principle. There are some valid uses, like when you are actually graphing 3 dimensions.
But 2d graphs, like line charts, bar charts, pie charts, that are given a 3rd dimension to make them look cool, all violate the principle of proportional ink. This is because projection effects, of representing 3d in 2d, will of necessity violate the principle.
This book was published apparently in August, 2020. Where was the LLM/Chatbot bubble at that time? This is not something I was tracking closely. ChatGPT was 2022, yes?
Searching for "language model" in the book finds nothing. Same for "LLM". The terms used to refer to big data models in this book appear to be:
"deep learning"
"machine learning"
The chapter starts with an excerpt from a July 8, 1958 New York Times article:
The Navy revealed the embryo of an electronic computer today that it expects will be able to walk, talk, see, write, reproduce itself and be conscious of its existence.
The embryo was called the "perceptron", invented by psychologist Frank Rosenblatt [1928–1971], "sometimes called the father of deep learning for his pioneering work on artificial neural networks." It looked suspiciously like a model of a neuron. "He described his work in grandiose terms" - thus setting a pattern that still lives on: the overhyping of neural network software.
In my 1st rant about the Bullshit Apocalypse, I showed pictures of the manuals & floppy disks I still have for "BrainMaker Software, Copyright 1988, 1989, 1990, Neural Network Simulation Software".
I had no idea the concept of neural networks dates all the way back to 1958. But, LOL, as our authors put it:
The same old magic is still selling tickets.
The authors jump forward to December 28, 2013, for another breathless NYT article.
It was as if the original writers outslept Rip Van Winkle, waking fifty-five years later to write the same article about the same technology using the same superlatives.
It still is all neural networks.
Most of the recent breakthroughs in machine learning—a subdiscipline of AI that studies algorithms designed to learn from data—can be ascribed to enormous leaps in the amount of data available and the processing power to deal with it, rather than to a fundamentally different approach.
I seem to have read that the Transformer algorithm, developed in 2017 at Google, also played a big part in the current LLM craze. I think it allowed parallel processing to greatly improve the training time for LLMs. Originally the algorithm was used for translation, then people realized it could be used for text generation, and, hello talking dogs! ChatGPT came out in 2022.
The authors introduce us to the building blocks of machine learning: "training data", "labels", "test data".
The authors seem to be building a case questioning the training data used in machine learning. We are reminded of an old computer acronym:
The authors discuss the bullshit hype that seems to be a constant of this technology. Drafting a robot bill of rights, worrying about whether robots & AI will take over and/or destroy the human race: this is just hype, hot air to inflate the actual, real world capabilities of these systems - hot air to inflate the bubble. It's bullshit, nothing more.
They talk about an article I remember, from 2017 - really, that long ago? - about some Facebook chatbot project where the bots started developing their very own (stupid) language. Horrors! They can plot against us! Facebook terminated the chatbots, not because they were afraid, but because they weren't meeting the goal of human-like communication.
I guess lots of people do like to be afraid. Not me. I quit watching horror movies decades ago.
A discussion of machine vision introduces us to machine learning concepts and issues:
If the accuracy of the computer’s labeling drops significantly when it moves from training data to test data, the model is likely overfitting—classifying noise as relevant information when making predictions. Overfitting is the bane of machine learning.
...
The challenge in machine learning is to develop algorithms that can generalize, applying what they have learned to identify patterns they haven’t seen before.
LLMs still suck at generalization. Take them outside the domain or range of their training data and they perform poorly.
The section titled "Garbage In, Garbage Out" begins:
Why is it so important to understand the role of training data in machine learning? Because it is here where the most catastrophic mistakes are made—and it is here where an educated eye can call bullshit on machine learning applications.
Ha ha, I like this quote from computing pioneer Charles Babbage [1791-1871]:
No algorithm, no matter how logically sound, can overcome flawed training data. Early computing pioneer Charles Babbage commented on this in the nineteenth century: “On two occasions I have been asked, ‘Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?’…I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.”
The authors move from machine vision to something more topical. But, this really seems dated. I don't think they realized how bad it would get, and how quickly.
Compare the challenge of teaching an algorithm to classify news stories as true or fake. ... You don’t necessarily know the answer just by looking at it; you may have to do some research. It’s not clear where to do that research, or what sources count as authoritative. Once you find an answer, reasonable people still might not agree whether you are looking at fake news, hyperpartisan news, satire, or some other type of misinformation. And because fake news is continually evolving, a training set of fake news stories from 2020 might be out of date by 2021.
Clearly they did not recognize the impact of the Bullshit Apocalypse, or "Slopocalypse Now" as Gary Marcus calls it. I'd venture that at this point, the slop has already won.
To call bullshit on the newest AI startup, it is often sufficient to ask for details about the training data. Where did it come from? Who labeled it? How representative is the data?
Where? Everywhere. Who? No one. Representative? Who knows.
Next we have a section "Gaydar Machines and Bullshit Conclusions".
In early September 2017, The Economist and The Guardian released a pair of oddly credulous news stories claiming that researchers at Stanford had developed AI capable of detecting a person’s sexual orientation.
And the machine was better at it than humans were. Well, duh, the machine was trained on 15,000 pictures from a dating site, the humans weren't!
The researchers based their conclusion on 2 theories, but then really gave no detailed data to support those theories. The machine as black box was not giving up any secrets. Our authors invoke "extraordinary claims require extraordinary evidence" and call bullshit.
The penultimate section of this chapter is titled "How Machines Think". Answer: we don't know. They're black boxes.
This opacity is a major challenge for the field of machine learning. The core purpose of the technology is to save people having to tell the computer what to learn in order to achieve its objective. Instead, the machines create their own rules to make decisions—and these rules often make little sense to humans.
[This recent post of mine talks about a Gary Marcus post that gives some insight into these machine rules, and where they go wrong: semantic leakage, subliminal learning, weird generalization, & inductive backdoors.]
They discuss the now well-known phenomenon of "machine indoctrination": where current biases, phobias, and other human cognitive errors are replicated in machines via their training data, leading to "algorithmic biases".
This next is, I guess, good news [my bold]:
To address the major impact that algorithms have on human lives, researchers and policy makers alike have started to call for algorithmic accountability and algorithmic transparency. Algorithmic accountability is the principle that firms or agencies using algorithms to make decisions are still responsible for those decisions, especially decisions that involve humans. We cannot let people excuse unjust or harmful actions by saying “It wasn’t our decision; it was the algorithm that did that.” Algorithmic transparency is the principle that people affected by decision-making algorithms should have a right to know why the algorithms are making the choices that they do. But many algorithms are considered trade secrets.
A converse of "algorithmic responsibility" seems to be the decision early in 2025 to not allow copyrights on purely AI-generated works. Holding a copyright is a human-only right/privilege.
"Algorithmic transparency" is going to be a tough nut to crack. They're black boxes, remember?
This next I think is still relevant:
Philosophers often describe knowledge as “justified true belief.” To know something, it has to be true and we have to believe it—but we also have to be able to provide a justification for our belief. That is not the case with our new machine companions. They have a different way of knowing and can offer a valuable complement to our own powers. But to be maximally useful, we often need to know how and why they make the decisions that they do. Society will need to decide where, and with which kinds of problems, this tradeoff—opacity for efficiency—is acceptable.
But, it's also dated. We can't worry about these niceties! We've got to beat China!
The final section of this chapter is titled "How Machines Fail". The early success but eventual failure of "Google Flu" - using the presence of flu-related keywords in Google queries to track & predict the spread of flu - is discussed. I think this was an early case of "AI slop poisoning" - autocomplete basically got into feedback loops & the downward spiral that comes when LLMs eat their own dogfood happened quickly.
The authors accurately predict some of the coming problems, with the machines getting enough quality data. I'm really interested in their read on where we are now. I think that their conclusion to this chapter is still valid.
In 2014, TED Conferences and the XPrize Foundation announced an award for “the first artificial intelligence to come to this stage and give a TED Talk compelling enough to win a standing ovation from the audience.” People worry that AI has surpassed humans, but we doubt AI will claim this award anytime soon. One might think that the TED brand of bullshit is just a cocktail of sound-bite science, management-speak, and techno-optimism. But it’s not so easy. You have to stir these elements together just right, and you have to sound like you believe them. For the foreseeable future, computers won’t be able to make the grade. Humans are far better bullshitters.
This statement begins the chapter, followed by descriptions of some of the awesome things physics (not sure why other disciplines weren't also touted) has discovered. But ...
For all that it does, however, it would be a mistake to conclude that science, as practiced today, provides an unerring conduit to the heart of ultimate reality. Rather, science is a haphazard collection of institutions, norms, customs, and traditions that have developed by trial and error over the past several centuries.
So this chapter is going to be about the institutions and culture of modern science: universities, research departments, professors, (grad) students, grants, peer review, "publish or perish", "the priority rule" - being the 1st to publish a discovery, ...
I think I might have started with "The scientific method is humanity's greatest invention." Fewer "howevers" to that statement. Organized science has lots of bureaurocracy involved, which means I don't think there's going to be any problem finding bullshit.
One of the reasons that science works so well is that it is self-correcting. Every claim is open to challenge and every fact or model can be overturned in the face of evidence. But while this organized skepticism makes science perhaps the best-known methodology for cutting through bullshit, it is not an absolute guarantor against such. There’s plenty of bullshit in science, some accidental and some deliberate.
Well-known "discoveries" that didn't pan out are listed. There is a discussion of what motivates scientists. They are humans, after all. We learn about "epistemically pure, and epistemically sullied." - science for the love of knowledge & discovery, or, to make a living, to get status.
Scientists write papers on experimental results or theories. Papers involving experiments should include enough detail that other scientists can try to reproduce the results. These go through the peer review process, and hopefully are published in journals, of varying prestige.
Our 1st inkling that all is not well:
Around the turn of the twenty-first century, replication problems began to crop up in a number of fields at unexpectedly high rates.
To understand what's up with this, we're going to have to learn some more statistical concepts, in particular the p-value.
Loosely speaking, a p-value tells us how likely it is that the pattern we’ve seen could have arisen by chance alone. If that is highly unlikely, we say that the result is statistically significant.
In what is, again, an excellent example, we learn about "the prosecutor's fallacy". The defendant's fingerprint, pulled from the FBI database of 50,000,000 fingerprints, matches the crime scene. The chance of a false match is 1 in 10,000,000. This implies 5 false matches, while there is only 1 correct match - only 1 guilty person. So the chance that the defendant is guilty based on the fingerprint match is only 1/6. I just looked it up in the wikipedia article on p-value, it says that the default for "predefined threshold value α which is referred to as the alpha level or significance level", is usually 0.05. 1/6 is 0.16, well above, so, no significance to the match.
Once again, we run into conditional probabilities, written as P(a|b), meaning "the probability that a is true, given that b is true". In the fallacy example, "P(match|innocent) = 1/10,000,000 whereas P(innocent|match) = 5/6."
A 2nd example is about 2 kids trying to discover ESP by guessing cards. We learn that the hypothesis that the results are explained by chance alone is called the "null hypothesis"; the hypothesis that the results are significant is called the "alternative hypothesis".
Getting back to science:
So here’s the dirty secret about p-values in science. When scientists report p-values, they’re doing something a bit like the prosecutor did in reporting the chance of an innocent person matching the fingerprint from the crime scene. They would like to know the probability that their null hypothesis is wrong, in light of the data they have observed. But that’s not what a p-value is. A p-value describes the probability of getting data at least as extreme as those observed, if the null hypothesis were true. ... Scientists are stuck using p-values because they don’t have a good way to calculate the probability of the alternative hypothesis.
I think we're getting to the crux of it here:
The point here is that the probability of the alternative hypothesis after we see the data, P(H1 | data), depends on the probability of the alternative hypothesis before we see the data, and that’s something that isn’t readily measured and incorporated into a scientific paper. So as scientists we do what we can do instead of what we want to do. We report P(data at least as extreme as we observed | H0), and this is what we call a p-value.
So what does all of this have to do with bullshit? Well, even scientists sometimes get confused about what p-values mean. Moreover, as scientific results pass from the scientific literature into press releases, newspapers, magazines, television programs, and so forth, p-values are often described inaccurately.
We've made it to a section titled "P-hacking and Publication Bias". OK, they confirm 0.05 p-value as the normal cutoff for significance.
Researchers are more interested in reading about statistically significant “positive” results than nonsignificant “negative” results, so both authors and journals have strong incentives to present significant results.
So the pressure is on, to get that significant result, to get that low p-value.
Very few scientists would commit scientific fraud to get the p-values they want, but there are numerous gray areas that still undermine the integrity of the scientific process. Researchers sometimes try different statistical assumptions or tests until they find a way to nudge their p-values across that critical p = 0.05 threshold of statistical significance. This is known as p-hacking, and it’s a serious problem.
The way to avoid this bad behavior:
Whatever it is that I choose to look at, the important thing is that I specify it clearly before I analyze the data. Otherwise by looking at enough different hypotheses, I’ll always find some significant results, even if there are no real patterns.
This sounds like "data dredging", which we learned about in Chapter 4, yes?
They demonstrate how easily it is to be tempted. You have all this data, years of work, p-value > 0.05. But if you use a subset of the data - say only one gender in a medical study - you can get the p-value below 0.05.
Congratulations. You’ve just p-hacked your study.
Is this next true? It doesn't seem right to me ...
Imaging a thousand researchers of unimpeachable integrity, all of whom refuse to p-hack under any circumstances. These virtuous scholars test a thousand hypotheses about relationships between political victories and analgesic use, all of which are false. Simply by chance, roughly fifty of these hypotheses will be statistically supported at the p = 0.05 level.
I think this next is the secret sauce.
There was no relationship. The appearance of a connection is purely an artifact of what kinds of results are considered worth publishing.
The fundamental problem here is that the chance that a paper gets published is not independent of the p-value that it reports. As a consequence, we slam head-on into a selection bias problem. The set of published papers represents a biased sample of the set of all experiments conducted. Significant results are strongly overrepresented in the literature, and nonsignificant results are underrepresented. The data from experiments that generated nonsignificant results end up in scientists’ file cabinets (or file systems, these days). This is what is sometimes called the file drawer effect.
OK, I think it makes sense. Favoring only significant p-values introduces selection bias. Wow, selection bias is really sneaky.
And, damn, Goodhart's Law, from Chapter 5, has also been triggered. Using p-value as a target has ruined its original purpose.
But the web gets still more tangled, as we learn about the "base rate fallacy": if you are testing for something that is very rare - Lyme disease in the example - false positives will overwhelm true positives, such that the test is worthless. Here's the big reveal:
In his paper “Why Most Published Research Findings Are False,” [epidemiologist John] Ioannidis draws the analogy between scientific studies and the interpretation of medical tests. He assumes that because of publication bias, most negative findings go unpublished and the literature comprises mostly positive results. If scientists are testing improbable hypotheses, the majority of positive results will be false positives, just as the majority of tests for Lyme disease, absent other risk factors, will be false positives.
Our authors aren't completely on board.
Ioannidis is overly pessimistic because he makes unrealistic assumptions about the kinds of hypotheses that researchers decide to test.
Interesting, the FDA requires all clinical trials of drugs to be reported, whether positive or not. So it's a great data set for studying publication bias.
Just as a sailor sees only the portion of the iceberg above the water’s surface, a researcher reads only the positive results in the scientific literature. This makes it difficult to tell how many negative results are lying beneath the waterline.
An iceberg, indeed!
The next section is titled "Clickbait Science". It discusses how information on scientific research makes its way to the general public. They give several examples of the divide between what scientists believe vs everybody else, on topics like GMO foods, climate change, evolution. Standard culture war stuff.
Some misformation about these topics is spread deliberately: big tobacco, big oil, fundamentalist religious groups all have agendas.
Still, a fair share of the blame lies squarely on the shoulders of scientists and science reporters.
Publication bias is even worse in mainstream media. LOL, they reference one of my fav riffs: "Yeah, I read where kale is good for you - this week."
One week, a single daily glass of wine increases heart disease risk, and the next week that same glass of wine decreases this risk.
The section "The Market for Bullshit Science" looks at the scientific publishing industry. We see Goodhart's Law again: scientists being measured by how many articles they published lead to the development of "predatory publishers", who will publish almost anything if you pay the $$$.
The authors examine peer review, and emphasize that it is not a panacea: "any scientific paper can be wrong", and "peer review does not guarantee that published papers are correct".
They give some suggestions on how to judge the quality of a paper.
The chapter ends with the section "Why Science Works".
Finally, irrespective of the problems we have discussed in this chapter, science just plain works. As we stated at the start of this chapter, science allows us to understand the nature of the physical world at scales far beyond what our senses evolved to detect and our minds evolved to comprehend. Equipped with this understanding, we have been able to create technologies that would seem magical to those only a few generations prior. Empirically, science is successful. Individual papers may be wrong and individual studies misreported in the popular press, but the institution as a whole is strong. We should keep this in perspective when we compare science to much of the other human knowledge—and human bullshit—that is out there.
In the preceding chapters, we have learned a lot about bullshit - and statistics, and dataviz, and institutional science. The last 2 chapters will put it all together. They both read like "how to" manuals. We're going to be given simple instructions on how to put our new-found knowledge to work, detecting, calling out, and, bonus, refuting bullshit.
While developing a rigorous bullshit detector is a lifelong project, one can go a long way with a few simple tricks that we will introduce in this chapter.
I'm going to maintain their numbering of our detection toolkit.
QUESTION THE SOURCE OF INFORMATION
Journalists are trained to ask the following simple questions about any piece of information they encounter:
Who is telling me this?
How does he or she know it?
What is this person trying to sell me?
The "who" is probably the easiest. Checking credentials is always a good idea.
"How they know" the authors say is harder.
They elaborate on the 3rd question:
Everyone is trying to sell you something; it is just a matter of figuring out what.
Is that really true? Not just being cynical? But "follow the money" is always good advice.
BEWARE OF UNFAIR COMPARISONS
It's funny how folk wisdom keeps popping up. We all know not to compare "apples and oranges".
The authors give an example of a list of "most dangerous cities". It is heavily influenced by whether or not the suburbs were counted as part of the city.
IF IT SEEMS TOO GOOD OR TOO BAD TO BE TRUE…
The example here is from early 2017, when, after new Agent Orange travel & immigration restrictions, NBC News tweeted:
“International student applications are down nearly 40 percent, survey shows.”
This set off the authors' bullshit detectors - the timing was all wrong. They dug down & found the NBC News report that the tweet referenced, which stated:
Applications from international students ... are down this year at nearly 40 percent of schools that answered a recent survey by the American Association of Collegiate Registrars and Admissions Officers.
Wow, what a complete misreprentation! And it's worse than that:
Yes, international applications decreased at 39 percent of universities—but they increased at 35 percent of universities. Taken together this isn’t news, it’s statistical noise.
An interesting characterization of this issue?
Here is an example of how a statement can be true and still qualify as bullshit.
THINK IN ORDERS OF MAGNITUDE
LOL, I was trained as a physicist, you don't have to tell me that. Physicists live & breathe orders of magnitude.
When people use bullshit numbers to support their arguments, they are often so far off that we can spot the bullshit by intuition and refute it without much research.
They give several examples, where some quick mental calculations immediately reveal bullshit.
When my 4 kids (3 x 800 SAT math score) were growing up, at the dinner table, someone would call "let's do the math!", & we'd spitball #s in our heads. But my kids all told me that doing mental math amazed most other people. I bet it's way worse now. I think we can conclude that innumeracy is definitely a friend to bullshit.
AVOID CONFIRMATION BIAS
This is, I think, a tall order. LOL, I like this quote:
Our susceptibility to confirmation bias can be seen as falling under the umbrella of sociologist Neil Postman’s dictum, “At any given time, the chief source of bullshit with which you have to contend is yourself.”
Confirmation bias is also a significant contributor to the spread of misinformation on the Internet. Why fact-check something you “know” is true?
The example they give is really kind of scary. The assertion is very believable, but, again, when you dig down, someone completely misinterpreted something.
CONSIDER MULTIPLE HYPOTHESES
In this chapter we have mainly looked how you can spot bullshit in the form of incorrect facts. But bullshit also arises in the form of incorrect explanations for true statements. The key thing to realize is that just because someone has an explanation for some phenomenon doesn’t mean that it is the explanation for that phenomenon.
In the example they give, there is an explanation that sounds plausible, but which, on further investigation, is completely false - LOL, it violates post hoc ergo propter hoc that we learned about in Chapter 4.
The chapter ends with a section titled "SPOTTING BULLSHIT ONLINE". They give 10 suggestions for this task. Most are, I think common sense. Here are a few that were new to me:
4. Use reverse image loopup. ... This is one of the more underutilized tools on the Web for fact-checking. If you are suspicious of a Twitter or Facebook account, check to see if the profile photo comes from a stock photo website.
I have never done this, I should check it out. A large percentage of new followers on BlueSky seem to look like the same slender Asian woman. These usually wind up as deleted accounts.
6. Take advantage of fact-checking organizations.
I know about these, I almost never use them. Lazy, I guess.
10. Reduce your information intake. Take a break; be bored a few times a day and revel in “missing out” instead of being anxious about what you missed. This will enhance your ability to process information with skepticism when you are online.
Interesting idea. Does it work?
I got off of social media, except for Twitter to announce song videos, in 2016. I got back on SubStack and BlueSky in late 2024, to count coup after the Blue Wave election result! Oops. SubStack has serious people and good content. Blogging and following RSS feeds I still think is best.
The section and chapter concludes with this exhortation. Interesting that they use a "littering" metaphor. Someone proposed fines for AI slop, analogous to fines for littering.
Most important: When you are using social media, remember the mantra “think more, share less.” The volume of information on social media, and the speed at which it allows us to interact, can be addictive. But as responsible citizens, we need to keep our information environments as clean as possible. Over the past half century people have learned not to litter the sides of the interstates. We need to do the same on the information superhighway. Online, we need to stop throwing our garbage out the car window and driving away into the anonymous night.
Hmmm, we're getting a bonus feature for our $$$: this chapter will go beyond just calling bullshit and discuss refuting it as well.
But 1st, the book will live up to its name. The authors define "calling bullshit"
Calling bullshit is a performative utterance in which one repudiates something objectionable. The scope of targets is broader than bullshit alone. You can call bullshit on bullshit, but you can also call bullshit on lies, treachery, trickery, or injustice.
I don't know that I've ever heard the term "performative utterance" before. The wikipedia article is not too too long, I'll give it a shot.
The authors say that statements can be declarative or imperative (what about interrogative?).
In the wryly titled book How to Do Things with Words, philosopher J. L. Austin[1911-1960]noted that there is yet a third class of things that we do with speech. There are utterances that, when we make them in the appropriate circumstances, are better viewed as actions than as expressions of a proposition. These are known as performative utterances. “I dub thee Knight”; “I christen this ship the HMS Beagle”; “I do [take this man as my lawfully wedded husband]”; “I do solemnly swear that I will support and defend the Constitution of the United States against all enemies, foreign and domestic.” In each case, the speaker is not merely reporting on her action, she is acting by means of her speech. Austin calls these sentences performative utterances, because one performs an action by uttering the expression.
...
The subject is usually “I,” and they are in the present tense rather than past or future tenses: “I resign” instead of “I resigned” or “I will resign.”
...
In addition to grammatical cues, the English language even has a somewhat archaic word, “hereby,” that can be used to flag a performative utterance if it is not obvious from context.
And, from the Wikipedia articla:
Performative utterances are not true or false, that is, not truth-evaluable; instead when something is wrong with them then they are "unhappy", while if nothing is wrong they are "happy".
Amazing the stuff that's out there that you've never heard of. Maybe this has some bearing on the nature of what LLMs are???
Calling bullshit is itself a performative utterance—and this observation is important for understanding what it means to call bullshit upon some claim. When I call bullshit, I am not merely reporting that I am skeptical of something you said. Rather, I am explicitly and often publicly pronouncing my disbelief. Why does this matter? Performative utterances are not idle talk. They are powerful acts, to be used with prudence. Calling bullshit is the same. Don’t call bullshit carelessly—but if you can, call bullshit when necessary.
So calling bullshit is some serious shit. Who knew? It makes sense though.
The authors discuss very seriously what serious business this is. Calling bullshit is a public thing. They give various guidelines, which I think can be summarized by the Bronze Rule: don't be an asshole.
This next is important.
We understand that the proper target of calling bullshit is an idea, not a person.
Calling bullshit, good; ad hominem attacks, bad.
The authors now go even 1 step further! From just calling bullshit to refuting bullshit! In the list that follows, each technique for refuting bullshit of course comes with wonderful & easy to understand examples.
Following Chapter 7, correcting misleading dataviz is a slam dunk argument.
Deploy a null model;
The point of a null model is not to accurately model the world, but rather to show that a pattern X, which has been interpreted as evidence of a process Y, could actually have arisen without Y occurring at all.
The next section is titled "The Psychology of Debunking". They make an excellent point here - is this identity politics? The term they use is cultural identity. They were talking to a zoo director about how to deal with PETA.
We explained how their views were entangled with their identities in a way that his were not. For example, if they were to convince him that keeping elephants in captivity was unethical, he would still be a scholar and zoo director. But if he were to persuade them that keeping elephants in zoos was justifiable, they could not retain their identities as PETA activists.
They provide 5 "time-tested tips for debunking myths":
Keep it simple.
Take it offline.
Find common ground.
Don't overemphasize the myth.
Fill in knowledge gaps with alternative explanations.
The authors now give their final formula for calling bullshit:
Be correct
Be charitable
Admit fault
- This 1 I can say is a strong point for me. The best organizations I worked in had very few people who were afraid to admit fault.
Be clear
Be pertinent
- Ouch! This one could also be "don't be a weiss-alles (know-it-all)". Working in tech, there are so many guys who think their mission in life is to be the smartest person in the room.
In the end, a well-actually guy [weiss-alles] has more in common with a bullshitter than he does with a caller of bullshit. A bullshitter disregards truth or logical coherence in order to impress or overwhelm an audience. That’s the well-actually guy. He doesn’t care about advancing truth, or about the logical coherence of his objections. He is simply trying to impress or intimidate someone with his knowledge. Calling bullshit is not about making yourself look or feel smarter. If that is your goal, you are missing the point of this chapter, and indeed the whole point of this book. Effective bullshit calling is about making others smarter. This should be your barometer of success—and requires an extra level of social finesse.
Well, that brings us to the end! Here's the last 3 paragraphs of the book. These remind me of the word "wisdom".
Above all, remember Neil Postman’s dictum: “At any given time, the chief source of bullshit with which you have to contend is yourself.” Confirmation bias can make us more confident than we ought to be, and humility is an important corrective. Self-reflection and an appreciation for the difficulty of getting to the truth: These are the marks of a mature thinker worth trusting. Sure, we want to keep the rest of the world honest—but for everyone’s sake, let’s start with ourselves.
Calling bullshit is more than a party trick, a confidence booster, or a way to sound impressive in front of your boss. It is a moral imperative. As we note in the opening line of the book, the world is awash with bullshit—from clickbait to deepfakes. Some of it is innocuous, some is a minor annoyance, and some is even funny. But a lot of the bullshit out there has serious consequences for human health and prosperity, the integrity of science, and democratic decision making.
The rise of misinformation and disinformation keeps us up at night. No law or fancy new AI is going to solve the problem. We all have to be a little more vigilant, a little more thoughtful, a little more careful when sharing information—and every once in a while, we need to call bullshit when we see it.
I really learned a lot from this book. It is really well organized; explanations are so clear; and the examples given are uniformly great. I'm sure these 2 college professors get great evaluations from their students - I presume if they can write this clearly that they are also great in-person instructors.
Kobo says it takes 7-8 hours to read. I think it is worth most people's time. I strongly recommend it to you.
And the ebook is only $4.99! I bought 2 trade paperback copies for friends, I think they were $20 each.
I am so curious as to their read on the current state of the Bullshit Apocalypse. They did such a good job on this book, the state of bullshit in 2020. You know, I think they could do a 2nd edition, and just basically update Chapter 8 on Big Data, and probably Chapters 10 & 11, where they tie together all the prior chapters.
I guess I am going to have to take their online course. I did the 1st lesson, I skipped the LLM hands-on, I don't think it hampered me from understanding the lesson. Hopefully that will continue to be the case in the rest of the lessons.
Note, in addition to the Bullshit & AI tags in this blog, I have tagged this post with Cognitive Illusions, aka cognitive error or cognitive biases.