No View Is the Whole: The World and the Ways We Make Sense of It (Complete Text)
Fifteen chapters in four parts, seventeen figures, seventy-five sources you can check. From a horse in Berlin that seemed to do arithmetic and the 1854 London cholera map to a programme alarm on the lunar lander, one question runs through it: how the way we make sense of the world is itself formed, checked, and handed on to the next person.
This is the complete text of the book, the same words as the single-page bilingual reader. The reader switches between Chinese and English and opens every figure at full size; this is the site-native version, with a sidebar table of contents, dark mode, and comments. Revised edition of 15 September 2026: four parts, fifteen chapters.
Preface — Counting Yourself In
When someone disagrees with us, an explanation is never hard to find. He has not met certain kinds of people, he was once burned in some particular way, or he has spent his whole life inside that one circle. When it comes to ourselves, we may feel there is nothing to explain: things are simply as they are.
Pressed further, we can add a line: “I have thought about this carefully.” The sentence carries great weight with ourselves; with others it may not. When the other person offers the same assurance, we still want to know what he has seen, whom he trusts, and which situations he has missed. Psychological research has found that when people assess bias in themselves and in others, they give different weight to inner intentions and to outward behaviour.1 We know how to trace another person’s judgement back to its sources; when the trail reaches ourselves, it sometimes stops.
What I want to pursue is what happens if we take that step further. The self that is used to understand the world also lives in the world: it learnt its methods from certain people, trained its skills in a certain environment, kept paying attention because it cared about certain things, and only then arrived at the answer that now seems as plain as anything could be. That history of formation affects where the answer can be used.
I call the shift that comes with recognising this “the second realisation”. We already know how to explain why other people think as they do; now the way of seeing that is doing the explaining must submit to the same questioning. Even our most successful experiences and the methods of reasoning we trust most are included.
Success is especially good at hiding this question. When a skill has helped for a long time, we pay less and less attention to the conditions it depends on. Move somewhere new, and the same way of working suddenly fails; the first thought may still be that other people are not cooperating, or that we have not tried hard enough. Putting in more effort sometimes improves the result, and sometimes only applies a method that no longer fits more thoroughly than before.
A simple problem about finding a ball can make this easily overlooked difference visible. Someone puts a ball in a box and leaves the room; while he is away, the ball is moved into a drawer. A child who has watched the whole thing is asked where the man will look first when he comes back.
The child knows the ball is in the drawer. To expect the man to look in the box, the child has to count in one more fact: he did not see the ball being moved. Problems of this kind have been used to study children’s understanding of other people’s beliefs.2 The same room holds the ball’s actual position and also a person who does not yet know the ball has gone. The second is enough to explain why he walks towards a box we know full well is empty.
Adults usually have no trouble with this scene. But when we say “the data are clear enough” or “anyone with experience knows this”, we have often stepped back onto our own familiar ground. What knowledge does it take to see that clarity? Under what circumstances was that experience acquired? If someone else had the same information and the same history, would we be willing to trust him too?
Figure 0.1 Use compatible standards of checking for yourself and for others. The information available to each side may differ, and the difference needs to be explained; your own certainty is not thereby exempt from checking. Drawn for this book.
Seeing ourselves this way also changes what we ask of thinking. At work we know time is limited, yet in reflection we may demand that every matter be thought through before we allow ourselves to stop. How much time a piece of reasoning had, what data it obtained, whether a method worked out by earlier people was available: all of these affect what it can accomplish. Asking only whether we have thought deeply enough is not yet a complete understanding of what thinking is.
Sometimes it is even that feeling we cannot put into words that first notices a detail the analysis never obtained. The experienced hand feels the sound is wrong; the novice hears only that the machine is still running. Putting both in front of the same written description has not necessarily given them the same information. To know when that feeling deserves trust, we need to trace how it was learnt and how it was checked.
None of these questions can be handed to someone else for good. To consult an expert, we still have to know whom our question is suited to; to decide that liking something is enough, we have already judged that this matter allows a choice made on liking alone. Life keeps requiring us to choose how to judge, even though most of the time we do it naturally, without giving it a name.
Fortunately, this capacity need not be invented from scratch. Earlier people have already tried many ways of understanding evidence, responsibility, conflict and living together, and have left reasons that can be learnt and criticised. We can draw on them, then recognise where they need to be revised. The coherence that gradually builds up comes from having reasons to follow when choosing a method as well: the same demands cannot be used only to scrutinise others, and exceptions must arise from differences in circumstance, not from who happens to benefit.
This is the layer I hope the book will help readers see more of. Faced with a highly persuasive answer, one can understand what it says and also see how it came to be possible; faced with difficulty, one can doubt the answer and also recognise whether what is needed is a different way of asking, another kind of information, or other people’s participation.
Some questions only appear at that point. The original answer still runs smoothly, yet we begin to notice that the understanding we have been using all along has left no room for certain things.
Part One — What Belief Rests On
1 — How a Fact Becomes a Sentence
In the early years of the twentieth century, a horse in Berlin called Clever Hans drew people who came to watch him answer questions. His owner, Wilhelm von Osten, had been a schoolteacher and believed that animals could be educated. He would set an arithmetic problem, and Hans would tap with a forehoof, seeming to stop at the right count. The spectators counted the taps as they fell and saw him, with their own eyes, get the answer right time after time.
The explanation that comes to mind first is cheating: was the owner secretly giving a signal? Investigation ran into trouble here. When the owner was asked to step away and someone else put the questions, Hans could still answer correctly. Simply accusing the owner of fraud did not account for everything that was happening in front of people.
The psychologist Oskar Pfungst devised a special kind of addition problem. The owner would first lean close to Hans’s ear and whisper a number that no one else could hear; Pfungst would then say another number, keeping it from the owner in the same way. Hans was then asked to add the two together.
This time each of the two men knew only one number, and neither knew the sum. If Hans had understood both numbers and could do addition, he would have been the only one present who might know the answer. After each trial, the researchers repeated it with the answer known to them, to see whether the result changed.
Of the thirty-one trials Pfungst recorded in which no one knew the answer, Hans got only three right; of the thirty-one in which the questioner knew, he got twenty-nine. Other tests found that blocking his view of the questioner also affected his answers. These differences turned Pfungst’s attention towards the people. What exactly was the horse seeing in them?
A questioner would often lean slightly forward, watching the horse’s foot and waiting for the taps to reach the answer. At the expected tap, the head would lift very slightly and the body would gradually return to a more upright posture. For the person, this may have been nothing more than an unconscious easing as the wait came to an end; for Hans, it was the signal to stop tapping. When Pfungst deliberately altered these movements, he could also affect when the horse stopped.3
Hans really could stop at the right count. What needed re-examining was why people took this performance as a capacity for arithmetic. “He tapped five times” records a result; “he worked out that it was five” goes further and explains how he arrived at that result. The first sentence being true is not enough to prove the second.
What is in front of you has already been selected
What a video recording of a street keeps depends on where the camera points, how the light falls and how often a frame is captured. Human observation selects as well. On the same street, a driver may attend to the traffic, an architect to the fronts of the houses, while a child is drawn to a dog by the kerb. They face the same street, yet what they remember may differ.
Choosing what to notice and interpreting what has been seen can be separated further still. “He did not reply to my message today” can be checked against the message log; “he is distancing himself from me” needs more to be known. He may have been busy, may have missed it, or may in fact be pulling away. If the second sentence is taken straight away as something confirmed, the other causes are easily overlooked.
“He is distancing himself from me” may also be true. The difficulty is that once we believe it, we pay more attention to the replies that do not come and remember less often the times he got in touch first. The idea that began as an interpretation of the record starts choosing the next batch of records for us.
Can an observation be found that involves no interpretation at all?
And yet even “he did not reply to my message” requires the record to be read. We have to know what sending and replying mean, fix the period we are counting, and notice whether he responded by some other route. If observation too depends on concepts, is the distinction drawn above still of any use?
Kant placed this difficulty of knowing at the centre of his thought. In the Critique of Pure Reason, published in 1781, he argued that our experience is organised by the forms of sensibility and the categories of the understanding; we cannot remove all of these conditions of knowing and then compare what remains with a world never presented through them. He also tried to show that it is precisely because experience has these shared conditions that certain knowledge can possess necessity. Space, time, causation and geometry each have a role in this account.4
This pushes the problem further. If knowing always passes through human faculties and methods, how are we to judge that one claim is more reliable than another? The geometry of Chapter 7 will bring a new problem to one piece of knowledge on which Kant relied heavily.
For now, a more concrete comparison can be made. To check whether someone interrupted a speaker, a recording that can be replayed usually offers more detail than the sentence “his attitude was bad”. We can count the interruptions, hear clearly the sentences before and after, and point to where our memories do not agree. The recording still has the limits of its framing and its sound pickup, but those limits can be established one by one.
The recording an event leaves behind can be edited, described in words, and then understood as an act of contempt. That understanding in turn changes what is recorded next time and how questions are asked. We can follow these changes back to the disputed step, but it is hard to arrange them as a fixed staircase that leads only forward.
Making an accusation something that can be answered
In a high-conflict setting, repeatedly reminding everyone “not to be subjective” is almost useless. What helps more is changing the shape of the sentence.
Turn “he simply does not respect me” into: “He interrupted me twice just now, and he did not respond to the risk I raised; I read those two actions as disrespect, though I do not yet know what he was thinking at the time.”
Turn “customers have no need for this feature at all” into: “None of the seven people interviewed so far brought it up unprompted, which makes me lower my estimate of how strong the need is; but the sample and the way the questions were put may both have affected this result.”
A dispute that matters is worth a few more sentences to set out the behaviour seen and one’s own reading of it. Only then does the other person know which point they can answer: admit to having interrupted you, explain why they did not respond, or point to a stretch of conversation you failed to note. Everyday exchanges that are familiar and undisputed naturally need not be unfolded like this at every sentence.
Memory goes on writing too
In 1974, Elizabeth Loftus and John Palmer had participants watch films of car accidents and then estimate the speed of the vehicles. The questions used different verbs for the collision, and the stronger wording produced higher estimates of speed. The researchers then pressed a further question: were participants simply adjusting the number upward to fit the wording as they answered, or had the memory itself been affected afterwards?
A second experiment recruited a fresh set of participants. After watching the film, some received the speed question with one wording or another, and some were not asked it at all. A week later, the researchers asked them whether they had seen any broken glass. There was no broken glass in the film, yet the group that had earlier met the stronger collision word more often answered that they had.5
The wording had already changed the later answers. The speaker need not have meant to mislead, and the person remembering need not have felt they were adding anything, yet broken glass appeared in an account that originally had none.
The researchers could still replay the film and know which answers did not match the picture. Everyday conflicts rarely have so complete a point of comparison: we rely on memory to say what happened, and other people’s accounts take part in how we remember afterwards. We begin to need records, and to need other people; and both come with selections of their own.
2 — You Do Not Have to Know Everything Yourself
In 1854, cholera broke out in Soho, London. The disease can cause severe diarrhoea and dehydration, and in the neighbourhood people died one after another. The influential view of the day linked disease with foul air: stinking surroundings often went together with sickness, and for people at the time the explanation was not without appeal.
The physician John Snow suspected instead that an important route of transmission was contaminated drinking water. He had already put forward claims to that effect before this outbreak; now he needed to find out whether what was happening in this district could lend them support. He looked into where the dead had lived, where their water came from and how they had lived, and he also drew the data onto a map.
Look at the map first. Around Broad Street, short black bars are stacked thickly along the edges of the houses. Each mark points to one death, and the positions labelled PUMP, the water pumps, are scattered among the streets.
Figure 2.1 Snow’s map of the deaths, from his book, recording the Broad Street cholera outbreak of 1854. Short black bars mark deaths; the clusters on the map provided leads for the investigation. Original: John Snow, public domain; see the source for the digital file.6
On the map the deaths cluster around the Broad Street pump. But does living close by mean having drunk its water? Some people drank water at home, some at their place of work, and some may have had it brought from a distance by someone else. To connect an address with a source of drinking water, each person’s circumstances still had to be looked into.
Snow’s record preserves a few places that especially invite further questions. In the surrounding streets people died one after another, yet the workhouse had relatively few deaths; the brewery, not far from the pump, had no deaths registered among its workers either. Why were the people in these places different?
The workhouse had 535 inmates already living there, and five of them died of cholera; it had its own well and its own supply, and sent nobody to Broad Street for water. For the brewery, Snow went and asked the proprietor, Mr Huggins. He said the workers had an allowance of malt liquor to drink. He believed they did not drink water at all, but he was clear about one thing he could be certain of: they did not take water from the pump in the street. The brewery also had a deep well and other supplies of its own.
The opposite kind of clue came from far away. A woman living in Hampstead died of cholera, though she had not been near Broad Street for several months. Only after Snow asked her son did he learn that she liked the water there, and that a cart regularly brought a large bottle of it to her house. A niece who came to visit drank it too, and after returning to her own home she fell ill and died. The two fell ill and died in places far apart, yet the water they had drunk came from one source.
On the evening of 7 September, Snow took his drinking-water inquiry to the local governing body; the next day the pump handle was removed. Before it was removed, the outbreak had already been declining and many residents had already left. The later fall in deaths therefore cannot tell us on its own how much the removal of the handle achieved. To assess the explanation of transmission by drinking water, we still have to look at the places that were close by yet suffered few deaths, the people who lived far away yet drank the same water, and the comparison between different sources of supply.7
How many people a map depends on
During the same epidemic, Snow was also comparing the customers of two water companies in south London. The two companies’ pipes ran through the same streets, and sometimes next-door neighbours took water from different companies; the Lambeth company had moved its intake upstream to cleaner water, while the intake of the Southwark and Vauxhall company was still affected by the city’s sewage. People who lived very close together, in similar conditions, might therefore be drinking water from different sources.
To make the comparison, he first had to find out which company supplied each household. Tenants did not necessarily know, since the water rate was paid by the landlord, so after asking the tenant the search had to go on; some people had been moved elsewhere after falling ill, and the record had to be traced back to their original address. Snow set the deaths in the first seven weeks of the epidemic against the number of houses each company supplied. The death rate per house for the Southwark and Vauxhall company came to roughly eight to nine times that of the Lambeth company.7
This inquiry had to check what residents and landlords said against the original addresses of the dead and the records of supply before it could compare deaths among households under different conditions of supply. Snow could not learn everything simply by walking into a street himself; the memories of residents, the papers landlords kept and the official records each answered part of the question. Anyone who later wanted to check his conclusion also had to know how he confirmed addresses, told the sources of supply apart and counted the deaths in each group.
Because Snow wrote down how the inquiry was carried out, we can examine his reasoning many years later without interviewing those residents again. Research teams often divide their work in the same way: some people measure, some organise the data, some analyse; and the instruments and mathematical methods each of them uses carry the results of earlier researchers inside them. If everyone had to invent their own instruments and rebuild all knowledge personally before beginning their own research, far less would get done.
Knowledge left behind in this way already exceeds any one person’s memory. The residents knew only about their own households, the water companies remembered their pipes, and Snow placed the two in a single inquiry. Their knowledge did not have to become the same first; together they could still answer a question that nobody had been able to answer before.
We can still read today about the cases that do not fit the first impression easily: the works close by with few deaths, the woman far away who drank water from the same source. Snow kept them, so that those who accepted his conclusion could go on asking questions. If all we remembered was one brave physician who stood against the majority, the most useful part of this knowledge would be the part that never got handed down.
An engineer who has handled bridge structures well many times gives us reason to value his judgement on such matters. That record also rests on measurements, colleagues and the original conditions on site; when he is brought to a new site, that support does not necessarily come with him. As for his opinions on education policy, the old record offers even less reason. What needs checking there is a different kind of knowledge, and how he obtains and compares evidence in that field.
How much a reader can check also depends on how the material is presented. If a report sets out its methods, the limits of its data and its later corrections, others can redo the calculations or question the conclusion. A screenshot with author, date and context cut away may leave us unable even to confirm what the original words were answering. Even when both speak with equal certainty, we have reason to trust them to different degrees.
Sometimes the cheapest effective next step is precisely to recover the source that was left out. To check a quotation, find the original first; to assess a method, first see what kind of problem it handles and where it fails. Guessing on your own that ‘this ought to make sense’ may take more time and still never touch the material that is actually missing.
Which part, exactly, is unknown
‘I don’t know’ does not always admit the same kind of lack. Return to Snow’s inquiry: the death records could be quite definite, while which water a particular household actually drank still needed a visit. Once the source of supply was established, whether the two were causally related required in turn a comparison with other households and with the possible explanations. What is already known does not all become void because questions remain.
Stating your doubt specifically makes it easier to decide the next step. If you do not know whether something happened, you can look for records or witnesses; if the cause is unclear, you can compare what several explanations would predict; if the facts are largely clear but a certain cost is unacceptable to you, then the trade-off has to be discussed, rather than the same batch of data checked over and over.
Suppose you are considering a job in another town. The salary can be confirmed with the company, the actual commute can be tried out, but whether you are willing to spend less time with your family cannot be settled by another salary report. Keeping the three questions apart avoids gathering data without end, and it also stops you from calling everything a matter of personal values when some of it can still be checked.
You can even go to the interview and try the commute first, without yet agreeing to move. Action brings back the information that was missing. By the time the decision really has to be made, the salary and the post may still be attractive, while the commute you have actually measured has changed how you see the job.
Experts who wait for time to check them
Listening to an expert’s analysis, what you feel most immediately is how fluently he speaks. A prediction, though, requires waiting. Whether a certain regime changes, whether an economic indicator crosses a threshold, will not hand in its answer when the programme ends.
From the mid-1980s, Philip Tetlock tracked the predictions of experts in politics and economics over a long period. He asked participants to make checkable judgements about specific events and to state their degree of confidence, then set these against the later outcomes and against comparison baselines.8
Performance in the study varied by person, by question and by method of assessment, and some did worse than simple comparison methods. Expertise that could previously be felt only in conversation now had a record that accumulated over time. Judging a person no longer had to start afresh from the impression left each time he spoke.
The record has to keep the prediction’s deadline, its outcome and the confidence held at the time, or there is nothing to compare later. A prophecy with no time limit and no explicit conditions may be described as not yet wrong for a very long time; those waiting on it have no way of knowing at what point it actually deserves their trust.
Putting trust where it fits
A record of predictions can reveal another difference that is easily confused: a person whose stated probabilities match the actual frequencies still may not help us tell which situations are more likely to happen. A set of hypothetical weather forecasts shows the difference.
Suppose it rains on fifty days out of a hundred. The first forecaster gives a fifty per cent chance every day. The second gives twenty per cent on fifty of the days and eighty per cent on the other fifty; it turns out to rain on ten days in the first group and forty in the second.
Both forecasters’ confidence matches the frequency of rain, but the second supplies one more piece of information useful to anyone going out: which days are more likely to be wet. This simplified data separates calibration from discrimination. Calibration deals with whether the stated confidence has a matching frequency; users usually also care how far the different days can actually be told apart.
To check whether ‘eighty per cent sure’ is reliable, we at least have to say clearly what is being predicted, what counts as its happening, and which outcomes it will be checked against. In everyday exchanges where estimating probabilities does not suit, we can still state our grounds: ‘This part I have done myself, that part rests on a report, and for the new situation there is no data yet.’ The listener then knows which claims come from experience and which still need separate confirmation.
There is also a limit that no amount of extra effort removes. After checking the expert, you can go on to check the people who assess experts, and then the institutions those people belong to. Every check draws on other knowledge. This road never suddenly delivers you to a position where you need trust nobody at all.
When you find that several reports were all copied from the same source, what looked like the agreement of many people loses some of its weight. Discovering that a prediction never left a checkable deadline has a similar effect. We can check only some of these things, but such specific findings are already enough to change our trust, without waiting until the whole body of knowledge has been examined.
Nor can the responsibility for checking be pushed entirely onto the user. If only the service provider can obtain the original records, yet the user is required to prove for themselves which step of the system went wrong, many problems can never be raised. The provider should give an intelligible account, suitable data for checking and a channel for handling disputes; otherwise ‘check it yourself’ merely asks people to complete a task without the information it requires.
Knowledge left for the next person
When a judgement passes into someone else’s hands, it can arrive as a bare conclusion, or it can be left together with its data, its methods and the questions still unresolved. The second kind of handing on takes more trouble, but it makes it possible for the next person to discover that the answer no longer suits a new situation.
This also changes how we picture relying on others. In accepting a piece of research we do more than borrow the little that the researcher knows beyond us; we also connect ourselves to the records, instruments, methods and subsequent corrections. Individuals forget, leave, and even refuse to admit mistakes, yet the material they leave behind may still let others carry the work on.
A book carries this responsibility too. Readers cannot redo every study on the author’s behalf, so the author must give the important claims their sources and make the key steps of the reasoning findable. Trust does not release the author from giving that account. It is worth placing, often, precisely because others can still ask questions after it has been placed.
3 — Who Decides What Is Worth Looking At
In Lewis Carroll’s Alice’s Adventures in Wonderland, Alice and a crowd of animals climb out of a pool of tears, soaked through. Everyone needs to get dry, and the Dodo proposes a race.
The course is roughly a circle, though the shape does not matter. There is no starting signal for the field; whoever wants to run runs, and whoever wants to stop stops. After a while the Dodo declares the race over, and everyone crowds round to ask who has won.
It thinks for a long time and decides that everybody has won, and that everybody must have a prize. Who is to provide the prizes? The Dodo points at Alice. She hands round the sweets from her pocket, and there is exactly one each. But she is to have a prize too, so she brings out a thimble that was already hers. The Dodo solemnly presents it back to her, and everyone cheers. Alice finds the whole thing absurd, and takes the prize all the same.9
By the end of the run, everyone is dry. If all the Dodo has to show is that running like this helps to dry a body, it has a result to report. Should Alice press it on why the prizes all had to come from her, pointing once more at the dried feathers would be an answer to a different question.
How much “it achieved its purpose” can say in defence of an arrangement depends on what we are evaluating. An examination mark can help a teacher see whether a student is ready for the next course; using it to decide who deserves respect calls for reasons to be given separately. A mark that has been calculated without error cannot, on its own, show that the second use is justified.
We can set the question down in a park. First decide what is to be surveyed: shade in summer, maintenance costs, or whether a wheelchair can get through. The purpose settles which tools the surveyor picks up, where they linger and when they come back for a second look. One of the things expertise does is help people recognise the differences that work of this kind needs to see.
Then the facts begin to constrain the answers. A gradient has its own way of being measured, shade has its hours, and the upkeep of a material cannot be filled in to suit a position. Caring about different things makes for different surveys; once two people are answering the same question, it is still possible to compare which measurement is the more reliable.
A difficulty of another kind surfaces only when the manager has to divide limited space and a limited budget. Keeping more trees, widening the paths and holding down maintenance may not all reach their best at once. However precise the gradient data, it will not decide on everyone’s behalf who should bear a little more of the inconvenience.
If an overall score multiplies shade, access and cost each by a weight, then whoever sets the weights is shaping which needs are met first. The formula may compute very exactly; the choice of weights still has to be explained to the people it affects. On the other side, if the gradient really was measured wrongly, the figure should be corrected. Whether a measurement is accurate, and how a limited budget should be shared out, are two different disputes.
Stating the use first also lets us judge, in concrete terms, whether an omission is a problem. A running route map that leaves out rest stops may still be enough for planning distances; if the same map is used to guide wheelchair users, the stairs along the way must be marked. The strongest ground for criticising such a map is to point to the omission that gets in the way of what it claims to help with.
Which differences are worth keeping
When we organise information for a given use, we usually leave some details out. A financial statement does not record every conversation in the office, and a route map does not mark the position of every window. Whether a detail is kept depends on whether it would affect the judgements the user has to make.
A model can be used to record the features we care about and the relations between them. Recording a cup as a capacity, a material and a degree of heat resistance, for instance, helps someone judge whether it is fit for hot water; if the task is to arrange shipping, weight, dimensions and fragility matter more. Four uses can be compared here:
| Purpose | Kept | Set aside for now | Possible error |
|---|---|---|---|
| Drinking | Capacity, heat resistance, safety | Exact shape | Overlooking heat resistance |
| Logistics | Dimensions, weight, fragility | Feel in the hand | Overlooking breakage |
| Design | Proportion, texture, manufacturing process | Some logistics details | Looks displacing use |
| Forensics | Residue, fingerprints, timing | Whether it is pleasant to drink from | Destroying evidence |
Each of these sets of data is useful, and they cannot be swapped about at will. Knowing that a cup withstands heat does not tell the shipper how large a box is needed; the person lifting fingerprints may have to finish recording them before someone else washes the cup. By the time the inquiry begins, the purpose has already shaped what is kept.
What the work is meant to get us
The task in front of us usually has a further reason standing behind it. A customer service department that wants shorter calls may want them so that people waiting to be answered wait a little less. If staff transfer complicated problems away quickly, the call figures improve while customers queue again and again and retell their story; the service has not improved because of it.
When the call ends, the customer may join another queue for another line. The original record stops at the moment the phone is put down, but his problem carries on. To know whether the service has improved, the record has to follow him a stretch further.
Time also sets different purposes against one another. A project may skip necessary maintenance in order to launch on schedule; a department may push pending problems into next quarter in order to hit this quarter’s target. The progress in hand still has value, but the assessment has to count in the cost that has been deferred.
Abstraction has more than one direction
To class a cup as a “container” is to attend to the fact that it holds things; to class it as a “fragile item” is to attend to the risk of a knock. The two classifications keep different features, and there is no need to rank one above the other first.
The distinction also helps in assessing records that have been reduced to numbers. If an employee’s performance score counts only the cases he has closed himself, it cannot show how much time he spent training colleagues; sorting test results into normal and abnormal may no longer show how near a value came to the threshold; drawing a rectangular box round a pedestrian in an image keeps mainly position and size. When the next step is to judge long-term contribution, track how a value changes or understand where the pedestrian intends to go, other records may be needed.
Whether an omission causes a problem still depends on the later use. Some representations mainly rearrange information so that calculation or lookup becomes easier, and do not necessarily delete any of the original content. Chapter 4 compares this case using road networks and numerals. Judging whether a representation is good means actually looking at what it preserves, what it makes convenient, and what the present question needs.
Purposes are changed by understanding too
If a purpose could be fixed once and for all, the difficulties that followed would mostly be a matter of finding the means. But people often find that, having learnt a few things, what they want to accomplish has changed as well.
Someone who first understood caring as doing everything for the other person, and who then heard that person’s own account, begins to value leaving him his choices. At that point the new understanding has changed what “doing it well” means. Pursuing the old goal more efficiently might, if anything, intrude on him more deeply.
In learning a craft, this kind of change is especially hard to explain in advance. A beginner may want only to get the job done fast, and only later comes to pick out details he could not hear or see before. Those details make him willing to slow down, even to stop being satisfied with work he was once proud of. Had he been asked at the outset whether he wanted to put in all those hours, he would not necessarily have known what he would be putting them in for.
We learn because we value something, and what we learn alters the reasons we valued it. Purpose takes part in inquiry, and inquiry takes part in forming purpose. Draw the two as sharply sequenced steps and this whole stretch of experience has nowhere to go.
A change of wishes is itself worth looking back on. Learning may lead a person to a new interest; advertising, rewards and group pressure may also set him chasing things he never cared about before. One way to evaluate such a change is to ask whether he had the chance to encounter other options, to understand the cost, and to refuse when he no longer wished to go on.
Public decisions carry one further practical requirement. A wheelchair user points out that the ramp cannot be used, and the information has certainly entered the room; if the designer can still pass over it because it is not among the established metrics, the intervention has had no effect on the judgement. The value of more viewpoints has to be judged by whether they can change the scope of the problem and what is done about it afterwards.
Institutions run into difficulty at this same point. An institution needs a goal before it can begin to survey and to allocate resources, yet the survey may bring back lived experience capable of changing the goal. What began as a count of how many people could walk through the park turns up the fact that some people cannot get in at all; if the original scoring is still used to show that everything is in order, the added understanding becomes a marginal note that cannot touch the decision.
We cannot guarantee that every wish becomes better for being understood. But an arrangement that permits only the improvement of means, and never permits the question “what does doing it well actually mean” to be asked afresh, has already set an end point to what people may learn.
Part Two — Models, Their Uses, and Where They Apply
4 — A Different Representation Makes Thinking Possible
Look first at the centre of each of these two maps, then follow one line out to the suburbs.
Figure 4.1 Two historical maps from different decades. On the left, the Underground map of 1908, marked in the archive as public domain; on the right, the second edition of Beck’s 1933 map, image source and credit: David Rumsey Map Collection, David Rumsey Map Center, Stanford Libraries, CC BY-NC-SA 3.0. The two maps also carry the differences of a network that changed over the years.1011
In the left-hand map the railway still clings to a city in which streets, riverbanks and parks can be recognised. Suburban distances stretch the lines out, while the centre is crammed with station names and bends. The right-hand map keeps the connections between stations, straightens the lines into regular directions and makes room for the names. The London you see has been stretched, compressed and rearranged.
When Harry Beck put the design forward in 1931, the publicity department turned it down. Could passengers really read a map that strayed so far from geography? By 1933 the design was at last printed as a pocket folder, and demand brought further printings.12
A passenger holding the map mainly wants to know which line to take, which stations it passes in order, and where to change. Beck opened out the crowded centre and shortened the suburban gaps so that station names and interchanges were easier to pick out. Distances on the map therefore no longer follow geographic scale, while the connections between stations and their sequence still have to be accurate.
When you ask instead how long it takes to walk between two stations, the spacing on the right-hand map is no longer enough to answer. The purpose has changed, and so has the information that needs to be put back. An omission has concrete gains and losses, and these can be compared.
A visitor arriving in London for the first time may not yet be familiar even with the line colours and the interchange symbol. He needs a key to tell him how to read the map, and then station names to confirm his direction. A map that someone who knows the network can take in at a glance is not necessarily as simple for him. As for the engineers who maintain the track, they also need actual positions and equipment data, and the pocket travel map does not supply these.
How long an explanation should be therefore also depends on what the reader has already learnt. Cut the key that a newcomer needs and the page is cleaner, but he has to go about asking people what the symbols mean. A concise explanation should spare people irrelevant work while keeping the explanation needed to finish the task in hand.
A digital map can show the travel route first and let people tap open exit and walking information; a printed one can separate the main map, the key and supplementary pages. Users ordinarily read only the part they need, and when a new question arises they can still find further explanation.
Compare on the same network first
The two historical maps are twenty-five years apart, and the lines themselves were added to and removed. To see on its own what the redrawing brings, it is best to hold the connections fixed. The figure below therefore constructs a separate one-way network of eight places: in both drawings, A to H, the arrow directions and the links are exactly the same, and only the placement of the coordinates changes.
Figure 4.2 Two drawings of the same set of links. This is a hypothetical example drawn for this book, not the London network. The right-hand drawing is arranged by fewest steps: A; B, C; D, E; F, G; H.
What is the smallest number of steps from A to H? Each arrow counts as one step, and you may travel only in the direction of the arrow. If you pick any route on the left-hand drawing and follow it to the end, what you get may be only one route among several, and you still have to compare the others before you can confirm it is the shortest. The right-hand drawing searches in a different order: first list every place reachable in one step, then every place reachable in two, and work outwards layer by layer. When a place turns up a second time, you already know that the earlier route reached it in no more steps than this one, so there is no need to start again from here.
Starting from A, one step reaches B and C, two steps reach D and E, three steps reach F and G, and only the fourth step finds H. The search has already listed every place reachable in fewer than four steps, and H is not among them, so four steps is at once one route that has been found and the smallest number of steps required.
This method is called breadth-first search. It supplies an order of searching, and it also supplies the reason for being sure the answer is the shortest. As the network grows, following the same method and recording the places already reached and the steps taken avoids trying complete routes over and over. What learning an algorithm that others have worked out saves is exactly this kind of repeated fumbling.
Now suppose some links take one minute and others ten, and we want the quickest route. Taking one step fewer no longer necessarily takes less time; three links of ten minutes each may be slower than five links of one minute each. Breadth-first search can still find the route with the fewest steps, but to find the route that takes the least time, the time of every link has to enter the calculation, and a method suited to differing costs has to be used.
Write it differently and the calculation changes
One hundred and five is written 105. The zero in the middle seems to stand for nothing, yet it keeps the units apart from the hundreds. Write 15 and the two digits that carry value are still there, but the quantity has changed. Place-value notation lets the same symbol stand for different magnitudes in different positions, and it lets addition, subtraction, multiplication and division proceed step by step along those positions.
Write twenty-three times fourteen as 23 × 14 and you can split fourteen into ten and four, work out two hundred and thirty first, then ninety-two, and put them together to make three hundred and twenty-two. Long multiplication sets these relations out on paper, so that all the intermediate results need not be held in the head.
The Roman numerals XXIII and XIV also stand for twenty-three and fourteen. The quantities have not changed, but the decimal long multiplication above cannot conveniently be applied to them directly. The person calculating can switch to a tool such as an abacus, or first rewrite the numbers in a notation suited to the operation. Recording a quantity and being able to work out a product easily are two different requirements.
When the Underground map was redrawn, some geographic detail was omitted. The numerical example shows another possibility: the quantity is kept unchanged, and certain operations still become easier because of the representation. When Larkin and Simon discussed diagrams and reasoning in 1987, they distinguished two things: that the content of two representations can be derived from each other does not mean that finding a given answer takes the same amount of work.13
A diagram can present relations such as adjacency, crossing and sequence directly on the page, easing the reader’s burden of cross-checking in memory while reading. But two objects drawn close together do not necessarily have a causal relation, and an arrow may indicate nothing more than order. When using a diagram, it is still necessary to say what the points and lines stand for, lest the drawing hint in addition at conclusions that have not yet been established.
An arrangement on paper has a further effect that is easily overlooked: someone else can carry on from it. Long multiplication leaves its intermediate results, and another person can see which column carried wrongly; a network with its links drawn lets someone who took no part in the original discussion still look for another route. A representation preserves relations that people can operate on, and thinking can therefore be interrupted, handed over and resumed.
We often think of tools as things picked up only after the thinking is done. Here the order is not so tidy. Only after learning how to lay quantities out on paper can people reliably complete calculations that were hard to complete in the head; only after drawing out the dependencies among tasks does it become possible to notice a wait that had never been spoken of. A new representation also takes part in forming the ability.
What can be carried elsewhere
Roads, the passing of messages and dependencies between tasks can all be represented with points and lines. Before carrying the same method of calculation across, the meaning of the points and lines has to be confirmed: is a point a place, a recipient or a task? Does a line mean that one can pass, that something can be sent, or that something must be finished first? If costs are being calculated, it also has to be made clear which costs increase as a line is traversed, and whether they may be counted more than once.
For example, when water flows through pipes, the amount entering and leaving a point can be calculated; when a message is forwarded, it may be copied to many people at once. That both sides use the word “flow” is not enough to justify carrying over the same conservation relation. Only when the differences are set out explicitly can one see which step of the original calculation needs changing.
Borrowing a story needs the same comparison. Someone who went for years without results and then succeeded can let readers feel how hard the waiting was. To use this to urge another project to hold on for one more year, one has to establish: what has the past investment accumulated? Which thing might one more year of waiting change? What signs show that results are drawing near? A failed project may also have a long history of investment, and length of time by itself cannot tell the two apart.
What a simple action needs behind it
A few taps on a phone and a food order is sent. The user does not have to contact the kitchen, arrange the order of deliveries or sort out delivery addresses; these are shared among the platform, the restaurant and the courier. The simple action lets him place the order without understanding all the details.
But if the delivery address cannot be made out, or the order status does not match what the restaurant received, the process that was hidden behind the screen becomes important. Finding where the error lies may require comparing the order record, the address and the delivery status. If these data were never kept, the user can press the same button a few more times and still get no answer.
The folder icon on a computer presents the operations of storing and organising files on the screen. The user can copy, move and open files without having to know each time where the data are actually held. A representation of this kind, made for people to operate, is called an interface.
Some concepts work in a similar way. To say that a company’s profit has risen lets us discuss revenue, costs and investment first, without immediately reading through every transaction. If, however, profit on the books rises while cash keeps falling, it becomes necessary to look into receivables, the timing of payments or other relevant details. The summary that was useful before cannot, on its own, answer the new question.
“He is very conservative” is also a summary. If this person suddenly supports a proposal that would greatly change the present state of affairs, rather than concluding straight away that he has contradicted himself, one can ask what exactly he is protecting: an existing process, a certain value, or the interests of a group of people? A word that is convenient for everyday description needs to be spelt out afresh when it meets a counter-example.
The boundary can be redrawn, yet the consequences remain real
Some decisions need a clear threshold. The law fixes the day on which a person comes of age, a monitoring system sets the value that triggers an alarm, a screening procedure lays down when further tests are arranged. People have to act at some moment, and so they divide a continuously changing situation into a few categories.
Take a hypothetical risk score: seventy or above is classed as high risk, and sixty-nine falls short. That one point can decide whether additional review is triggered, yet the actual risk does not necessarily jump between sixty-nine and seventy.
Choosing the threshold requires reference to the relation between risk and score, and consideration of the respective consequences of missing someone at high risk and of misjudging someone at low risk, as well as how much review capacity can be committed. The data limit which choices are well founded; the decision-maker still has to explain why the line is drawn here.
Once the threshold is adopted, what was only one point in a calculation changes how a particular person is treated. Discussion of whether the line is drawn reasonably therefore cannot stay inside the formula.
Classification changes how people are treated, yet the power of renaming has places it cannot reach.
When a traveller finds the way by the map, the actual streets test whether the route exists. Draw a short cut on the map and the wall will not let anyone through on that account; write a higher permitted load for a bridge and the bridge does not become any stronger because of the number.
We can choose how to describe, but whether a description is useful is still limited by the thing described. This book calls this situation “the constraint of reality”: some outcomes cannot be changed by changing the words alone; the understanding of the causes, or the actual practice, has to be adjusted.
This gives different models a place where they can be compared. When two maps give opposite directions for the same road, the discussion cannot be closed with “a difference of viewpoint”; when two maps guide travel by train and on foot respectively, each may be accurate. Only after confirming whether they answer the same question does one know which kind of difference to check.
When both maps can be used
A city can have a transit map, a relief map, a population map and a house-price map all at once. A company, too, can be understood through its balance sheet, its operating procedures and its division of labour. Each of these representations helps answer a different question, and which to choose depends on what needs to be known this time.
Faced with the same question, several maps sometimes correct one another, and sometimes each still carries its own cost. Someone in a hurry may accept a rough estimate of time, while someone studying the causes of congestion has to keep several more variables. Comparing how the maps predict new data, what they assume and what it costs to check them can help with the choice; no single requirement can always come first for every use.
Organisations also meet another situation: two analyses are both quite reliable, yet they cannot combine themselves into a decision. The safety assessment points to a danger, the revenue forecast looks favourable, and the arithmetic in neither is wrong. If the two are converted into a single overall score, someone has to decide how much revenue offsets how much danger; if it is ruled that a certain kind of danger is enough to halt the project, reasons have to be given for that too.
Drawing more maps means this discussion need not proceed in a vacuum. The lines and numbers on the maps, though, will not take on the choice for the people at the table.
5 — What Was Left Out Is Still at Work
On the evening of 31 May 2009, Air France flight AF447 left Rio de Janeiro for Paris with two hundred and twenty-eight people on board. By the early hours of 1 June the aircraft was cruising over the Atlantic. The captain had handed over and left the cockpit to rest, leaving the two co-pilots in their seats, one flying, the other monitoring and assisting.
The airspeed readings suddenly became inconsistent for a short time, and the autopilot disconnected. The investigation concluded that the probes used to obtain airspeed information had most likely been obstructed for a while by ice crystals. The co-pilot flying took control and pulled back on the sidestick, raising the nose; the aircraft began to climb and, as it did so, to lose speed.
The other co-pilot noticed that the aircraft was climbing and asked several times for it to descend. The pilot flying did make nose-down inputs, and the climb eased for a moment, but then the nose came up again and stayed up. As the climb went on, the stall warning sounded continuously; the altitude reached about thirty-eight thousand feet at one point, some three thousand feet above the original cruising level. The monitoring co-pilot called repeatedly for the captain to come back.
A stall, here, is a matter of the wings. The angle at which the wing meets the airflow becomes too steep, the flow begins to separate from the wing surface, and lift falls away. With its nose pointing upwards the aircraft can still be dropping fast. Holding the nose up does not end that state merely because it looks like flying upwards.
Although some of the airspeed readings had come back, the crew still failed to recognise the stall and recover from it. After the captain returned to the cockpit the airspeed readings became invalid again and the stall warning stopped; when the nose was briefly lowered and the readings became valid once more, the warning sounded again. Whether the warning was on or off did not correspond directly to whether the aircraft was now any safer. It had to be read together with whether the airspeed data was valid, what attitude the aircraft was in, and how it was descending. The three men did not arrive at a correct judgement in time. The aircraft went into the sea, and no one on board survived.14
The French accident investigators traced the event through the control records, the logic of the warnings, the training, and the way the crew worked together. In ordinary flight the systems handle a great deal on the pilots’ behalf; once something goes wrong, details that normally need no individual attention suddenly become part of a judgement that has to be made at once.
A person can have the controls back without having, at the same moment, the ability to understand the situation. “When the system meets something it cannot handle, it hands over to a human” sounds like a thorough arrangement: the machine does the routine work and the human keeps the final decision. When the handover actually comes, what is left to the human may be exactly the situation that is least familiar and leaves least time to think.
A simple action still rests on many things
Putting a file into a folder on your own computer and putting a file into a folder on a remote server can be the same drag on screen. The first may finish almost at once; the second has to wait for the transfer. If the connection drops part-way, no amount of resemblance to a local folder will keep that communication alive.
The user only drags an icon; the program takes care of storing and transmitting, which is why the folder view can be so simple. Whether the file arrives still depends on whether the disk can be written to, whether the connection holds, and whether the remote device responds. These details do not normally need to be shown one by one, yet when they fail they bear on the very same action.
In 2002 the programmer Joel Spolsky gave this kind of situation a name, the “leaky abstraction”, and offered a generalisation about engineering that has travelled widely since: any abstraction with real substance will, to some degree, expose the details it set out to hide.15
Take TCP, the set of rules for network transmission that lets applications work with a reliable, ordered stream of data, and that can cope with some packets going missing or arriving out of order. A connection may still fail, and there is no fixed guarantee of how long transmission will take.16 Someone working with remote files who loses the connection may need to check whether the transfer completed, wait for it to recover, or reconnect.
Delay and disconnection may have been set out in the protocol’s specification all along; the trouble sometimes lies in an interface that gives the user no reminder, and in a user who mistakes everyday convenience for a guarantee of success at any moment. Once we are clear about what the specification actually promises, we can tell whether the tool has failed to deliver or whether we expected a capability it never agreed to provide.
This book keeps the term “leaky abstraction” to name that relationship: some of the differences that a representation or interface left out or hid can still affect the outcomes the user cares about; once they come to matter, working only within the original representation may no longer be enough to understand or deal with what is in front of us.
Figure 5.1 Hiding the details does not cut the dependence on them. The arrows in the figure show relations of support and influence, not that every use will fail. Drawn for this book.
This kind of limit can show up on the very first use. A newly set-up remote folder, for example, can no longer supply files the way a local folder does once the connection drops. The loss of skill over years of use, documentation going out of date, the difficulty of replacing the tool: these are a separate set of risks. Even if none of them ever arises, a connection remains a condition for any remote operation.
Why some gaps cannot be filled from within the original representation
Suppose two sets of data survive only as their averages. One set was forty and sixty; the other was zero and one hundred. Both average fifty, and there is nothing wrong with that calculation.
Now someone asks whether either set contains a value below twenty. The first does not; the second does. If the average really is all that is left, there is no way to tell from it which set was which. However precisely the fifty is computed, however long it is analysed, the difference that has been lost will not grow back out of it.
Figure 5.2 The same summary can stand for situations that call for different answers. The numbers are hypothetical, chosen for the argument, not measured data. Drawn for this book.
The reason can be stated quite plainly. If a representation records two actual situations as the same content, and some question demands different answers for the two, then any fixed way of judging that relies on this representation alone cannot answer correctly in both cases. At the least, one piece of information that separates them has to be added, or it has to be admitted that for now they cannot be told apart.
When the task is only to compute the average of the two numbers, either summary is entirely sufficient. An abstraction can stay accurate on a well-defined question; leaving something out does not automatically amount to being wrong. The difficulty is that we so often take a representation built for one kind of question and go on to answer other questions with it.
So long as the omitted difference stays irrelevant to the question, we can safely spare ourselves the effort of handling it. Once the question changes, more computation may be no help at all. The difficulty at that point may not even be one the person using the representation can resolve: the original data may be kept somewhere else, or it may never have been kept.
Having a limit, and leaving the limit nowhere to be found
Consider two hypothetical systems for reporting faults. In both, the user presses “report fault” once. One of them also saves the relevant raw data from that moment and allows situations outside the existing categories to be written in; the other keeps only the result once it has been sorted into the existing categories. The screens are equally simple; what can be looked up afterwards is not.
| Point of comparison | Version A | Version B |
|---|---|---|
| Everyday operation | One press to report a fault | One press to report a fault |
| Background record | Keeps the relevant raw data; allows uncategorised situations to be added | Keeps only the result after sorting into existing categories |
| When something must be traced | Raw records and supplementary notes are available | Discarded differences cannot be recovered from the result |
If it later turns out that faults filed under one category in fact had different causes, Version A can go back to the saved data and see which differences the original classification failed to record; Version B, having discarded the relevant details, cannot reconstruct them from the categorised results alone. Ease of operation, and the keeping of data for catching errors later, can be handled separately by different parts of the design.
Choosing between them also means reckoning the cost of storing data and weighing privacy and use. There is no need to keep everything against every remote possibility. But once we have reason to expect a particular failure that matters, whether to leave an adequate path for tracing it becomes a practical choice.
In the same way, a checkout rule that supports a single currency can refuse other currencies outright, or it can quietly treat figures in different currencies as if they were in the same unit. The first tells the user that another method is needed; the second may cover the problem with an answer that looks normal. Every abstraction has limits, and the ways of handling those limits can differ enormously.
Once the data has been kept, someone still has to be able to read it and act on it. That ability, too, has to be maintained.
When a system has run smoothly for a long time, people usually check it less often and spend their time on other things. That saves effort, but it can also let certain abilities go rusty. Someone who rarely handles faults by hand, for example, may need more time to recognise the situation when suddenly asked to take over.
Writing about automation in 1983, Lisanne Bainbridge pointed to a contradiction in practice. Once automation has taken over routine operation, what remains for the human may be the rare and difficult abnormality; yet the operator, lacking daily practice, finds it hard to grasp the situation quickly when suddenly required to. Keeping a person at the last gate does not by itself guarantee that the person is still able to complete the handover.17
Keeping people able to take over is therefore work that has to go on continuously in ordinary times. The practice required, the status information that must be available, and the time needed to act all have costs. The more smoothly the automation runs, the more easily these investments come to look superfluous: if they are so seldom used, why keep paying for them? Only when the abnormality arrives does the cost that was saved reappear in another form.
A remote file can lose its connection on the very first use; proficiency in a rare operation can decline after long disuse. These are difficulties from different sources, and design has to face both. Knowing that an abstraction cannot take care of everything does not, on its own, tell us how much capacity to hold in reserve, day to day, for an exception that has not yet arrived.
More and more reports, and the war no clearer
In 1961 Robert McNamara arrived at the United States Department of Defense. He came from wartime statistical work and corporate management, and he valued decisions supported by quantities, costs, and comparisons. These methods could reveal differences that had previously been hard to set side by side, and they gave analysts an important place in the running of defence.
The Vietnam War is often described afterwards as a war lost because the only thing anyone looked at was the enemy body count. The historian Gregory Daddis’s research complicates that story. What was collected at the time went well beyond kill counts: weapons captured, local security, the state of control over territory, a great mass of data. Part of the problem lay precisely in there being too much of it, and too little of a consistent way of telling which numbers actually meant progress.18
Political legitimacy, local networks, popular attitudes, and the opponent’s mobilisation could all shape the course of the war. To use figures for casualties, captured weapons, or local control, one had to understand under what conditions each of them reflected strategic progress. Adding another batch of numbers, if their relation to the strategic aims still could not be spelt out, did not necessarily add to anyone’s understanding of how the war was going.
When those reports were also used to allocate resources and assign responsibility, looking into what lay outside the numbers mattered all the more. Could observations from people on the ground supplement the existing reports? When someone found that an indicator did not match the actual state of control, could that prompt decision-makers to reassess? Gaps in measurement can persist because the decision process accepts only certain numbers.
Those who find a way round also know something
Sometimes a user can deal with a tool’s limits without first working out the full principle behind them.
Suppose an image tool consistently fails to pick out the edges of objects in a certain kind of photograph. After a few comparisons the user finds that adjusting the contrast first gives a selection closer to what is wanted. He may have no idea how the selection algorithm computes its result, yet he has learnt a useful technique: on this kind of photograph, changing the input first improves what follows.
This can be called a local workaround. It solves the immediate difficulty first, and lets repeated use and comparison confirm its scope of use afterwards. If it stops working on another kind of photograph, the rule has to be narrowed or the cause looked into afresh.
Success after success, though, can make people forget the scope. What was first written down was “on this kind of photograph, adjust the contrast first; the result is better”. After a few handovers only “always raise the contrast before selecting” remains. Those who come later follow the rule without knowing that some photographs never needed the treatment, and may even be made worse by it. The technique has become a general rule, and what has vanished is the set of comparisons that originally supported it.
We rely on this kind of limited grasp with a great many tools. A writer does not have to understand how a typeface is rendered on screen before writing; a maintenance engineer may, from years of comparison, hear an abnormal sound first and only then ask someone to test for the cause. Whether such a competence is reliable should be judged by what the person recognises, under what conditions it works, and whether he can improve after a miss.
If the technique keeps failing, the user may have to ask someone to examine the inputs and the environment, or switch to another measurement to cross-check. That takes time; overhauling the whole classification takes far more, and after the change there will be new limits to meet. Sometimes the only option is to narrow the use, or even suspend it, and bear the loss of not being able to do that thing for a while.
Each of these responses has its own price, and no single action can be called “dealing with the leak” once and for all. Still, whoever takes over can at least be spared some wasted effort: if the original user leaves behind the kinds of photograph the technique applies to, the cases where it failed, and the original files, the successor need not start guessing again from the single line “always raise the contrast”.
Leaving a record of how something was used can also, at times, let those who come later ask questions that no one had asked before.
Some unknowns can already be put as definite questions. Was a particular error caused by a change of camera? One can compare old and new images, rerun the recognition, and eliminate possibilities one by one.
Harder are the cases where the system has not yet recorded the difference that caused the error. Suppose an engineer keeps only the recognition results, with no original images and no record of when the camera was changed. When he sees accuracy falling, he is short of several clues he might have used to propose and test a cause. The original report had no fields for them, and an error does not sprout its own explanation.
We cannot list every surprise in advance, but we can give new findings somewhere to be written down: allowing faults outside the existing categories to be entered, letting operators add what they saw on the spot, or keeping the raw data connected to important results. Once a new pattern is found, one can then decide whether to add a category, change the measurement, or test again.
Which data to keep, and who may add to it, still have to be weighed against cost and privacy. The aim of the design is to give problems no one has yet foreseen a chance to be noticed and investigated, rather than to require every system to hoard everything in advance.
How science differs from a temporary fix
A tool can stay in use on the strength of local techniques, and scientific theories, too, are frequently revised in the face of new phenomena. The difference between the two can be seen in what checks the revision must then submit to.
The photograph technique above claims only that it improves selection on one kind of input, and repeated comparison on that kind of photograph gives it partial support. To go further and explain how features of the image affect the algorithm, one would have to put forward expectations that can be checked on new photographs; and when others repeat the work with different data, they should be able to see corresponding results. As the claim grows, so does the evidence it requires.
New concepts in science need the same kind of checking. The discussion of dark matter concerns the problem of mass in phenomena such as the motion of galaxies and gravitational lensing; dark energy concerns the explanation of the universe’s accelerating expansion. Each is constrained by several sets of observations, and each still has open questions and ongoing research.19 To assess them, one compares how the whole explanation fits the different lines of evidence, and what further observable results it leads us to expect.
If every discordant result were met only with an added explanation that cannot be checked separately, any theory could be kept safe. What gives us reason to trust a theory more is its ability to anticipate new situations from its principles and to submit to checking against independent data.
Does the fact that knowledge is still limited mean that “ultimate truth must be unattainable”? That step cannot be taken from present shortcomings alone. We can establish that a given representation cannot answer a given question without thereby establishing the limits of every future method.
Scientific realists care whether a theory describes structure the world actually has; instrumentalists lay more stress on a theory’s use in organising experience and making predictions. They understand theories differently, yet they can still jointly check what assumptions a claim uses, what evidence it fits, and what kind of finding would demand its revision.
Some limits call for a different road
When a leak appears, the existing system does not necessarily have to be repaired until it can handle everything. Giving up a category, stopping a scoring scheme, or retiring a process may be better.
If a form keeps squeezing out important experience, the answer need not be to add fields without end; perhaps decisions of this kind need interviews and case-by-case judgement to take part as well. If a method shifts its maintenance costs onto its users, keeping it running is not necessarily a goal worth putting first.
Choosing how to handle a limit also involves how the costs are distributed. A form that saves managers time in review may make it hard for the people filling it in to describe their real difficulties; keeping an old tool in service may leave frontline workers patching things by hand again and again. In weighing whether to keep it, the burden on these people and the alternatives should be compared together.
A form kept originally to save trouble may leave another group of people forever making up for what it failed to record. Whether it is worth keeping can no longer be judged by how efficiently the form organises things.
6 — Useful, True, and Worthwhile
It was getting dark, and the little girl was still in the street selling matches. Andersen sets the story on the last day of the year. Snow was falling, and from every house came light and the smell of roast goose, while she walked outside barefoot. The oversized slippers she had set out in were lost when she dodged out of the way of a carriage, and all day nobody had bought a match from her or given her a single coin.
She dared not go home. Her father would beat her for bringing back no money, and home was hardly warmer, for the wind still came in through the cracks in the roof. At last she crept into a corner between two houses, her hands stiff with cold. One match, struck, might warm her fingers a little.
The flame caught. She held her hands towards the light, and it was as if she were sitting before a great iron stove. Her body grew warm, and she was stretching out her feet to warm them too when the match went out. The stove was gone, and all that remained in her hand was the burnt stub.
She struck another. Where the light fell, the wall turned as thin as gauze, and behind it stood a table laid with a cloth, a roast goose still steaming on it. The goose jumped down from its dish and waddled towards her, knife and fork and all; then the flame died, and there was nothing in front of her again but the thick, cold, damp wall.
With the third match she found herself sitting beneath a splendid Christmas tree. Many small lights burned on its branches, and she reached up to touch them, and the match went out. The lights of the tree rose higher and higher until they became the stars in the sky, and one of them fell, drawing a long streak of light behind it. She remembered what her grandmother had told her: when a star falls, a soul is going up to God. Her grandmother was the one person who had loved her, and she was dead.
When the next match flared, her grandmother stood there in the light. The little girl begged to be taken with her. She knew by now what happened when the flame went out: the stove had vanished, the goose had vanished, the Christmas tree had vanished, and this time she would not lose her grandmother too. In haste she struck the whole bundle of matches at once, to keep her grandmother there.
Her grandmother lifted her up in her arms, and together they rose to a place where there was no cold, no hunger and no sorrow, and came to God.
The next morning, people in the street found the girl leaning against the wall, frozen to death, a bundle of burnt-out matches beside her. They supposed she had only been trying to warm herself. They did not know what she had seen in the light.20
When an experience really does bring comfort, are we weighing that comfort, or the reality the person was actually standing in?
A report that seems to be about you alone
Suppose a personality profile reads: “You care a great deal about how others see you, yet at times you wish you did not have to be swayed by them.” Reading a sentence like that, you may feel it has caught a contradiction in you exactly. But can that feeling of recognition show that the test has really told you apart from other people?
Bertram Forer asked the students in his class to complete a test. Afterwards he handed each of them a personality description that appeared to have been written from their results, and asked them to rate how well it fitted. The students generally rated it highly.
Only then came the disclosure: they had all been given the same text. Forer’s paper of 1949 called the exercise a classroom demonstration. What he was challenging was the practice of validating a diagnostic tool by the assent of the people it was applied to. If one description can make many people each feel “this is exactly me”, then the feeling of being seen cannot on its own prove that the description has picked out anything individual.21
A sentence in a personality report may apply to one person; the same sentence may apply to most of the class. To learn whether a test distinguishes between individuals, you have to set different people’s results side by side. It is not enough to let each person read their own and then ask, “Does this sound like you?”
Forer’s question was whether a test could recognise individual differences, and the subjects’ sense that the description fitted was not enough to answer it. Yet when the thing under study is itself a feeling, such as pain or discomfort, how the person describes their own change becomes data that cannot be done without.
The word placebo usually brings to mind a patient who does not know that what they have been given contains no active drug. In 2010, Ted Kaptchuk and his colleagues studied a different situation: if participants are told plainly, does the effect still appear?
The condition they studied was irritable bowel syndrome, an illness whose troubles include abdominal pain and changes in bowel habit. Participants were divided into groups. One group knew that what they were taking was a placebo and were given a positive account of the treatment; the other received a similar degree of contact with the clinicians but not the treatment itself. Over three weeks, the placebo group improved more on some of the self-reported symptom measures.22
What the study compared was an intervention made up of an explanation and a schedule of pills, and the improvement showed up mainly in the symptoms participants reported themselves. They knew which group they were in, and their expectations and the way they answered may also have shaped the measurements. The result is therefore something to go on studying: why did this arrangement help with some symptoms, and which of its ingredients made the difference?
Why planets go backwards in the sky
Some successes push the question the other way. When a method predicts things not previously known, and its results agree again and again with new observations, we have reason to believe it. But how much of it should we believe? Astronomers long used a system that could compute the positions of the planets, and the account of how the heavenly bodies actually move was later very substantially rewritten.
Record a planet’s position against the background stars over many nights in a row and you will find that it does not always move in the same direction. At times it seems to stop, then to go backwards, then to resume its former course. Ancient astronomy had to explain this retrograde motion, and it also had to calculate when it would occur.
In the second century, Ptolemy of Alexandria set out a geocentric system of astronomy in the Almagest. One of its devices is called the epicycle: the planet moves round a small circle, while the centre of that small circle moves round a larger one. Seen from the earth, the two motions combine to trace a path that sometimes advances and sometimes retreats. The full model had further geometrical arrangements as well, fitted to the observations of the different bodies.
This tradition, with the corrections made to it afterwards, was used for a very long time to compute planetary positions, and it preserved a great body of observations and methods of calculation. Later astronomers were able to compare the old explanation with new ones precisely because these records were there to use.
If the earth also travels round the sun, retrograde motion can be understood in a different way. The earth and an outer planet each move forward, and their relative positions keep changing; when the earth, on the inside track, overtakes the outer planet, the direction in which we look towards it shifts backwards against the background of stars. The retreat seen in the sky can then be explained by the relative motion of observer and planet.
The work Copernicus published in 1543 proposed a heliocentric arrangement, but it still used circular motion, and its computational accuracy left room for improvement. Later, Kepler, working from Tycho Brahe’s observations, developed the elliptical orbit and the relations that go with it, and these were then tested against new tables and new observations. The explanation of celestial motion and the improvement of positional prediction passed through distinct stages of research.23
The computational achievements of the old method did not vanish with what came later. It left behind a stock of usable predictions, and also an account of motion waiting to be rewritten. Researchers had to carry on from this mixed inheritance: which calculations to keep, which assumptions to re-examine, and which phenomena remained unexplained.
Some things are worth doing in themselves
Knowing that a piece of information can be checked still leaves the decision of whether it is worth checking now. Time is limited, and two pieces of information, equally reliable, may bear very differently on the decision in front of you.
An engineer dealing with a fault can begin by confirming which settings were recently changed, when the error appears, and which operations reproduce it. Every one of these details might repay study, but if only half an hour remains, the first thing to get hold of is the information most likely to change what is done next.
An astronomical question with no bearing on today’s decisions may still be worth a lifetime of study. Its value does not have to be adjudicated by the half hour spent on the fault.
There are many reasons to pursue the truth, and the same is true of other activities. If the only question asked is whether something makes a person more comfortable, we miss the things people do knowing full well that they will be uncomfortable.
Attending a funeral may include comfort, but a person may also go in order to acknowledge a relationship, to remember the dead together with others, or to complete a farewell they have chosen to take on. Even if the day brings no comfort at all, these reasons do not necessarily disappear.
Consider another hypothetical question. A survey asks you to write down how much you would pay to protect an endangered species. Someone refuses to enter a figure, perhaps because they believe that whether a species deserves to survive should not be settled by how much people are willing to pay.
If every such refusal is recorded as zero, an objection to pricing may be written down as complete indifference. The survey needs to allow people to say why they left the box empty. Conservation in practice still faces the difficulty of limited resources and still has to argue about how to allocate them; but before allocating, one has first to understand correctly what it is that people value.
Reading a novel, a person may enjoy the rhythm of the language, worry over the fate of a character, or feel again some stretch of life that resists being put into words. A passage can fail to persuade us as an argument and still console, or make us notice a problem we had been unwilling to face. Nor does a novel have to teach a transferable method before it deserves the time spent reading it.
The same activity may also make promises about matters of fact. If a novel is taken as an accurate record of a period of history, its history needs checking; if a ritual charges a fee on the promise of certain cure, the evidence for that effect has to be examined. Respecting what reading or ritual means to people does not require exempting these additional promises from checking.
If we intend to use an example we have read to support a decision, we take on an added duty of explanation. Chapter 4 discussed how someone else’s plan succeeding after many years cannot by itself show that our own plan should continue. This does not erase the strength the story once gave; it is only that, when the time comes to commit the next year, reasons that bear on our own situation still have to be found.
When comfort is bound up with a way of life
Some beliefs lack sufficient evidence yet make people more hopeful and more willing to look after one another. Acknowledging the good a belief does can be kept separate from accepting its account of the world.
Take the funeral discussed earlier in this chapter. A family uses the ceremony to remember the dead and to express how much the relationship meant, and these actions have reasons anyone can understand. If someone also believes that the dead have found peace in another world, others can respect what has been entrusted to that belief without claiming to have confirmed that such a world exists.
If the same belief is then used to make major decisions on behalf of other people, the question changes. To claim that a certain message is enough to settle someone else’s treatment, finances or safety involves effects that can be checked and risks that others will bear. Personal conviction and the experience of being comforted cannot on their own supply the evidence such decisions need.
A religion usually contains, all at once, an account of the origin of the universe, rituals, the norms of a community, ethical practice and a sense of identity. In discussing a particular religion, unless one first says which claim is in question, it is easy for each side to say a great deal without ever answering the same point.
To criticise an origin story for lacking evidence, for example, is not necessarily to deny the value of believers caring for one another; and to affirm the support a community provides is not enough to show that the story really happened. Different claims can receive different assessments, without first passing a single verdict on the whole tradition.
Science, too, comprises many activities: measurement, model-building, collective review and technical application. Its achievements help to show how these methods arrive at reliable knowledge. As for what ought to be pursued and which costs are acceptable, scientific research can supply the relevant facts, and a discussion of values and responsibility is still needed.
Real life, however, does not always lay these things out separately. A community’s beliefs about the cosmos may be the very reason its members are willing to care for one another; a line of scientific research is often thought worth pursuing because people already value the lives it might change. Analysis can tell the different reasons apart, but the person living inside them may lose, all at the same time, a certainty, a circle of friends and the direction their life had.
When a practice leaves the community it came from, these relations change again. It may be taken up to answer questions that were never posed in the same way before.
The Four Noble Truths of the Buddhist tradition speak of suffering, the origin of suffering, the cessation of suffering, and the path that leads to that cessation. The arrangement lets the practitioner recognise the predicament, understand how it arose, and then commit to practice. When some of these exercises are carried into a modern hospital, the original religious discipline and the clinical research need to be accounted for separately.
In 1979, Jon Kabat-Zinn began work at a stress reduction clinic at the University of Massachusetts Medical Center, organising exercises in breathing, bodily awareness and attention to present experience into a clinical course. The approach later became widely known as mindfulness-based stress reduction. Once inside the hospital, the practice had to face concrete questions: which complaints it helps with, compared with what, how long the effect lasts, and who it may not suit.24
Clinical research can begin by comparing the effect of the practice on particular complaints, and participants need not accept the whole religious cosmology. How the exercises were selected from the original tradition and adapted, and whether its ethical requirements were retained, belong to another part of understanding this transfer. Kabat-Zinn’s own account of the course also includes the attitude of practice and its background.
Take a hypothetical company that uses attention exercises to help its staff cope with stress. Employees may learn useful methods, and the company still needs to check whether the workload is reasonable, which processes create a continuing burden, and who is able to change them. If every discomfort is put down to insufficient personal practice, working conditions that could have been improved may go unaddressed indefinitely.
This is why “useful” has to say for whom, and over what span of time. An explanation may set an employee’s mind at rest for the day while allowing an unreasonable workload to continue; a simple classification saves the caseworker a few minutes while the person wrongly classified may spend days on an appeal. Chapter 14 will go on to trace how institutions make some of these results far easier to see than others.
How the truth reaches a person
An important truth can cause pain. Learning that a relationship is over, or that a project of many years has failed, brings no immediate relief. But the choices the person makes next may depend on knowing these things.
There are also different ways of telling. A person may be given time to understand, to ask questions, and to get help with the consequences; or they may be forced to respond at once in front of everyone. The truth of the news is the same; the treatment they receive is very different.
Harder still, we sometimes genuinely do not know how much they can bear at this moment. Fearing their pain may come from care, and it may also lead us to keep deciding on their behalf what they are allowed to know. They need support, and they have their own choices to make; and their response may change what we had understood “support” to mean.
When the little girl struck the whole bundle of matches, the story let her keep her grandmother. The reader who knows the ending can no longer see that blaze as nothing more than a comfort.
7 — Answers Come with Conditions
What do the interior angles of a triangle add up to?
Figure 7.1 On the left, a triangle on the Euclidean plane with straight lines for sides; on the right, a triangle on an ideal sphere with short arcs of great circles for sides. The two answers are 180° and 270°. A mathematical construction. Drawn for this book.
You can walk the right-hand figure as a route. Start at the North Pole and follow a meridian down to the equator, turn ninety degrees, walk a quarter of the way round the equator, then turn ninety degrees again and head back to the North Pole. The two meridians meet at the pole, and the angle between them there is ninety degrees as well. Three angles, two hundred and seventy degrees in all.
A triangle on the plane has interior angles adding up to one hundred and eighty degrees, yet the spherical triangle on the right has three right angles. The difference lies in the space the figure sits in, and in which paths we take as its sides. The example on the sphere uses short arcs of great circles; it does not violate the theorem about straight-sided triangles on the plane.
A mathematical theorem states which axioms and definitions it adopts, and then proves its conclusion. To use it to describe an actual object, one must also confirm that those assumptions match the purpose well enough. When surveying the shape of a small plot of ground, for example, a flat approximation may already be sufficient; a route across a large stretch of the Earth’s surface calls for attention to curvature and to the precision required.
The one they spent two thousand years trying to prove
Euclid’s Elements organises geometrical reasoning by first setting out a number of starting points that may be adopted. The famous fifth postulate says that if a straight line crosses two other lines, and the interior angles on the same side add up to less than two right angles, then those two lines, extended on that side, will meet.
This statement is more roundabout than some of the other starting points. For a long time mathematicians hoped to prove it from the remaining assumptions, so that it would no longer have to be accepted on its own. In 1733 Saccheri tried to reach a contradiction by denying the relevant assumption and following where that led. He obtained a good many results that would later look close to non-Euclidean geometry, and in the end he still tried to rule that road out.
In the nineteenth century Lobachevsky and Bolyai openly developed systems that differed from Euclidean geometry. Riemann, in a lecture of 1854, went further and studied more general spaces and geometries. It gradually became possible to compare geometries built on different assumptions: what each could prove, and whether contradictions arose.25
The spherical example at the start of this chapter gives the difference between geometries a visible shape before anything else. A geodesic on a sphere is the path that runs locally straightest along the surface; the angles at which geodesics meet on the sphere obey a different relation from the angles between straight lines on the plane.
Kant once used geometry to illustrate necessary knowledge that is not obtained from individual experiences alone. The development of several geometries made a distinction more pressing: does a proof within one system also suffice to decide which geometry physical space adopts?4 Answering the second question also requires comparing theory with actual observation.
When general relativity later described gravity through the geometry of curved spacetime, it had to submit to the check of physical observation. Mathematical study of the different possibilities gave physics more representations to draw on; whether they are fit to describe the world is judged by other evidence.
The small print the rule leaves off
Everyday rules seldom come with their conditions printed alongside. “First come, first served” settles a great many disputes when the matter is buying a drink; used to allocate help that is urgent and important, it raises the questions of who has the chance to arrive first, who can afford to wait, and whose need cannot be put off. The wording of the rule has not changed at all, yet the reasons for choosing it may fall short.
How long a practice is used also bears on whether it is suitable. Postponing a quarrel for a while may let both sides cool down; postponing it every time may leave an important matter unanswered for a long while. An approximate calculation carries only a small error over a short span, and once the errors accumulate through repetition they may exceed what is tolerable. When judging the effect, the length of use has to be reported along with it.
If “it depends” stops there, it is still of limited help. More useful is to point out which difference would change the judgement: whether the urgency is the same, whether the costs of waiting are wildly unequal, whether different stretches of the road take the same time. Once the conditions are concrete, what to check next is also clearer.
Newtonian mechanics has not vanished from the world
Newtonian mechanics supports a great deal of reliable work. In describing motion at everyday scales, and in designing many machines and engineering structures, a suitable approximation reaches the precision required. When relativity arrived, this work was not all made obsolete.
Satellite positioning, by contrast, needs extremely precise comparisons of time. Orbital motion and gravitational conditions affect the rate of a satellite’s clock relative to a clock on the ground; if these differences are not properly handled, the positioning result suffers. For this reason the system has to incorporate the corresponding relativistic corrections.26
So when using an approximate method, one can first state the precision required and the known error, and then decide whether it is enough. Everyday engineering need not adopt the most elaborate model every time; when higher precision is needed, or a different scale is encountered, the comparison should be made afresh.
The person who fed it every day
In The Problems of Philosophy Russell leaves us a chicken. The person who brings it food every day arrives one day and, instead of feeding it, kills it.27
The ending is so short that it hardly gives anyone time to prepare. Every past feeding was real; had the chicken kept a complete record, it could even have kept one without a single error. But those records did not hand over the keeper’s plans for the future along with them.
Those who know the ending find it easy to laugh at the bird. Yet the question Russell leaves behind pursues the reader just as closely: why should what has kept happening in the past go on happening? To answer with regularities that have succeeded before still leaves one having to explain why those regularities will remain in force. Digging further back into the past for a few more successes does not supply this demand with a wholly independent guarantee.
Everyday predictions still differ in how reliable they are. To predict whether the water supply will be interrupted, for example, one can note that there has been water every day in the past, and one can also check the state of the equipment, the maintenance schedule and the source. This information cannot guarantee that nothing will ever go wrong, yet it can identify risks that go unnoticed when only the past record is consulted.
In the same way, a satisfaction survey that interviews only existing customers can tell you about the experience of those who replied. To draw conclusions about everyone who ever came into contact with the service, one also needs to know about those who left, those who did not reply, and those who never managed to get into the service at all. Enlarging the sample of customers of the same kind can make certain estimates steadier; it cannot, on its own, fill in the opinions of those who were never surveyed.
When the mechanism still has to be traced
Snow did not first acquire complete microscopic knowledge of how cholera spreads before comparing different sources of drinking water. A suitably designed study can provide causal evidence that an intervention affects a particular outcome while the full mechanism is still unclear.
Studying how the effect comes about, in turn, helps in judging what will happen once the setting changes. Knowing which factors the effect depends on lets researchers give priority to testing the differences that might alter it. Comparing actual effects and establishing the mechanism of action can therefore help each other.
An intervention may also produce several effects at once. One of them may be beneficial while another cancels it out, so that the overall result does not necessarily match what was expected. Having explained one pathway of action, then, one still needs to compare the important outcomes directly. The clinical trial in the next chapter will make this difference concrete.
Knowing part of the mechanism may still miss a new situation. Anomalies reported by users, the raw data, and a second kind of measurement sometimes let researchers discover a difference they had not thought of. Chapter 5 discussed how to preserve these ways of catching errors; now the known scope of use can also be recorded more concretely.
Writing down the scope of use
“This is a retail model” or “this is manufacturing experience” is still too broad. Image models that all serve to inspect the appearance of objects may perform very differently depending on the lens, the lighting, the material of the object, or the angle of the shot. A shared name does not guarantee that the factors affecting recognition are shared too.
Take a hypothetical model used to flag damaged packaging in product photographs, which are then passed to staff for review. It has been tested only under fixed lighting, with a specified lens, and on a few kinds of packaging; backlighting and other reflective materials the team has not yet confirmed. Writing down these circumstances says more about where the result came from than the bare line “ninety-five per cent accuracy”.
Later the camera is replaced, while the program stays at the same version. Most of the photographs the team spot-checks are well lit, and the average score looks good, yet misses on the few backlit photographs have increased. If only the total score is kept, it is easy to conclude that the new lens made no difference. Keeping the disputed images, and comparing misses and false alarms separately, is what allows the conditions under which performance worsened to be found, little by little.
At this point “scope of use” has acquired new content. The team had thought of the lens simply as a tool for obtaining photographs of the same kind; now they have to include how it combines with lighting and material in their account. The scope of use is not always something known before the study begins, waiting to be printed as a line of small print beside the product; it also takes shape slowly, through failure and investigation.
If a later version improves performance under backlighting, the old limitation should be updated as well. The record needs to show which version was tested in which environment, and what changed afterwards. Otherwise new users may vouch for a new use with an old score, or may go on avoiding a difficulty that has already been solved.
The conditions in a geometrical theorem must be stated before the proof can begin. The conditions of use for an empirical method, by contrast, often become clear while the work is under way. Demanding that every condition be written down at the outset looks careful, and may in practice demand that people know in advance what the research is about to discover.
When a group of models do not cast the same vote
Studying the future climate means handling processes that act on one another: the atmosphere, the oceans, clouds and ice. Models built by different research teams make different arrangements in certain details. CMIP, the Coupled Model Intercomparison Project, has teams run their simulations under a common experimental design, so that the results are better placed to be compared.
Setting the results side by side helps show which changes are more consistent and where the differences are larger. But the models are not fully independent of one another; they may share data, methods or code. Having many models does not mean that each additional one brings an equally independent ballot of evidence.
Some of the models in CMIP6 give a higher equilibrium climate sensitivity. This quantity asks, roughly: if the concentration of carbon dioxide doubles and the climate reaches a new equilibrium, by how much does the global mean surface temperature change? To assess it, one can also draw on historical warming, palaeoclimate, and research into the relevant physical processes, rather than counting votes among model outputs alone.
The IPCC assessment of 2021 brought together several lines of evidence and described the range of sensitivity with different degrees of uncertainty. The ranges corresponding to “likely” and “very likely” differ; to understand them, the numerical intervals and their accompanying statements of uncertainty have to be read together.28
Climate scenarios, for their part, are projections under specified conditions of development. The familiar RCP8.5 is a high radiative forcing pathway, and the number in its name relates to the level of radiative forcing in 2100 relative to pre-industrial times. It helps researchers ask how the climate might change if development follows this set of conditions; how likely the conditions themselves are to come about has to be assessed separately.
Our grip on the future may therefore come from models and from evidence outside the models constraining one another. When the results do not yet agree, the disagreement can also leave a direction for research: is some physical process poorly understood, or were different development scenarios adopted? Getting to the bottom of these questions often takes far longer than computing an average across all the models.
8 — Giving Error a Chance to Show
The aircraft came back with bullet holes in them. Ground crews could walk up to the fuselage and inspect where the damage lay; the aircraft that did not come back could not be inspected in the same way.
During the Second World War this was a practical difficulty in adding protection to aircraft. Armour is heavy, and no part of the airframe can be thickened without limit; someone had to decide where a limited weight was best placed.
The damaged airframes that did return looked like the most direct material for deciding where the armour should go. They had also already passed through one round of selection: having taken this damage, they were still able to fly home.
Abraham Wald, working in the Statistical Research Group at Columbia, wrote a set of studies in 1943 that used the number of sorties, the losses, and the damage data from surviving aircraft to estimate the vulnerability of the aircraft.29
Many holes in one part of the returning aircraft might mean that part was hit more often, or it might mean that an aircraft hit there could still return fairly easily. Conversely, few holes might mean fewer hits, or it might mean that an aircraft hit there rarely came back. To estimate where the vulnerable places were, assumptions about being hit and about surviving had to be combined before anything about the lost aircraft could be inferred from the data on those that returned.
The records taken on the apron could be entirely free of error and still fail to represent every aircraft that flew. The problem arose before the data came into being: some aircraft did not return, and so left no damage record of the same kind. However carefully the returning aircraft were inspected, that inspection could not fill the gap on its own.
Hiring records can run into a similar limit. A company can see how the people it hired went on to perform, but it does not know how the people it turned away would have done in the same posts with the same training. Judging a selection method by the results of the first group alone leaves out an important comparison.
Supplying the comparison with those who were not hired also involves real vacancies, real training, and real pay. Seeing from the statistics which piece of data is missing does not make the cost of obtaining it disappear.
A small number of appeals may mean the service is good, or it may mean that appealing is difficult, that appeals go nowhere, or that the most dissatisfied people have already left. If the system records only the appeals that were successfully submitted, counting the same channel again will still miss those who tried and failed. One can instead conduct interviews, observe people using the process, or check how many gave up part-way through submitting.
The number of sources needs checking in the same way. Three reports that all copy from the same set of data do not provide three independent pieces of evidence; three questionnaires with the same too-narrow options may all miss the same kind of experience together. Before adding data, first confirm which of the missing parts this round of collection can supply.
Sometimes the person raising an objection changes the research question as well. Where the question had been the completion rate among participants, they point out that a certain group was never eligible to take part; where it had been user satisfaction, they point out that some people were compelled to accept the service and could not leave it. Questions like these bring people outside the original sample into the evaluation, and only then does it become easier to judge whom an institution has served and whom it has missed.
After pushing a number down
Medical research has met a heavier version of the problem: if a number that looks bad is pushed down, will patients live longer?
A myocardial infarction damages the heart muscle, and some patients afterwards develop premature ventricular contractions, an early beat slipped into the rhythm of the heart. The condition is associated with a higher risk of death. Certain drugs can suppress this kind of arrhythmia, and so a conjecture worth testing arose: if it is suppressed, will fewer people die?
The clinical trial known as CAST assigned eligible patients to a drug group or a placebo group and went on comparing what happened to them. The preliminary report published in 1989 showed that the groups taking the two drugs under test, encainide and flecainide, had more deaths from arrhythmia or non-fatal cardiac arrests, and more deaths overall. The trial of these two drugs was therefore stopped early. The rhythm indicator that the treatment had been meant to improve did not bring the expected benefit in survival.30
The trial illustrates the difference between a correlation and the effect of an intervention. That a certain rhythm is associated with a higher risk of death, and that a drug can reduce that rhythm, is still not enough to prove that taking the drug lowers mortality. The drug may have other effects at the same time, and the important outcomes for patients must be compared directly.
Checking has its own costs. Some methods can alter one condition and compare; others can only make use of differences that arise naturally; arrangements that touch other people’s interests carry a responsibility for risk and for fairness. Errors that are serious, persistent, and hard to undo are usually worth a fuller check, while small errors that are easy to spot and repair can be handled more lightly.
To test an explanation, look first at where it expects a different result from the other explanations. Clever Hans in the first chapter is an example: that the horse could still count correctly when the questioner did not know the answer, and that it needed to pick up a person’s signal to stop, are two expectations that can be compared.
If any result at all can be described as supporting the same theory, it is hard to know what evidence an observation has added. Someone attacks another person, and this is explained as a certain desire; he restrains himself from attacking, and this is explained as the same desire being repressed. Without some separate way of recognising the desire and the repression, these two accounts on their own can absorb opposite behaviours alike.
When the same behaviour, whether it appears or not, can be gathered into the same explanation, the researcher still has to point to some observable difference. Otherwise all we know is that he is able to keep talking.
What they hoped to see on the day of the eclipse
In ordinary daylight the stars near the sun are drowned in its glare. A total eclipse offers a brief opportunity: with the moon covering the sun, observers can photograph the surrounding star field and set it against comparison plates taken at other times.
On 29 May 1919, expeditions organised from Britain went separately to the island of Príncipe off the coast of West Africa and to Sobral in Brazil. Arthur Eddington took part in the former. What the study set out to compare was how far the apparent positions of the stars shifted as their light passed near the sun; general relativity gave a quantitative expectation for this deflection.
The expeditions had to deal with weather, the quality of the plates, the instruments, and measurement error before they could compare whether the shift in the stars’ positions matched the expectation. The results published afterwards supported general relativity, and the evidence came from the analysis of the photographs and of the errors.31
When Karl Popper later looked back on the development of his thinking, he set great store by theories that take a risk in advance like this. Compared with the kind of explanation just described, which can add an account for opposite behaviours alike, the expected amount of deflection gave observation a chance to come into conflict with the theory. This was an important reason for his demand that empirical theories possess falsifiability.32
Checks of this kind are aimed at predictions about the empirical world. Mathematical proof, ethical reasons, and literary understanding each have their own ways of being assessed; of a theory that claims to predict what will be observed, we can ask that it first state its expectation clearly and then give observation the chance to disagree with it.
Yet the person who makes this demand also meets moments when his own judgement needs looking at again.
If ‘survival of the fittest’ is put only as ‘whatever survives was well adapted’, it does look like going round in a circle. Defining fitness by survival and then explaining survival by fitness has yet to deliver any content that would separate one result from another.
Popper at one time rated Darwinism a valuable metaphysical research programme and had reservations about its testability. In 1978 he revised that assessment in an article, acknowledging that the theory of natural selection has testable content.32
Evolutionary research in practice can measure traits and environment before the outcome appears. Peter and Rosemary Grant and their research team followed the ground finches of Daphne Major in the Galápagos over many years, recording the characteristics of individual birds and whether they survived. A drought in 1977 changed the food supply, and different beak shapes were related to the ability to use the seeds that remained. The researchers could therefore compare which individuals survived and how the characteristics of the population changed afterwards.33
Recording beak shape and food first and comparing survival afterwards makes it possible to check a concrete relationship, rather than simply calling the survivors ‘the fittest’ after the fact. If the result does not match expectation, researchers can go on to examine their assessment of the food, their measurement of the traits, and other factors bearing on survival.
One anomaly, several possibilities
When an observation fails to match expectation, the difference may come from the data, from the procedure, from an auxiliary assumption, or from the theory itself. An anomalous figure may come from a faulty instrument, for instance, and conditions assumed to be fixed may have changed during the test.
The tests on Clever Hans in the first chapter could close in on the cause because they were arranged so that one thing, whether a person knew the answer, varied in a way that could be compared. Had the questions, the place, and the manner of asking all changed at once, a wrong answer would have been hard to attribute to any one of them. Which conditions need to be held constant, and which further situations need comparing, depends on the cause one is trying to rule out this time.
The same difficulty appears in everyday claims. ‘Giving people more authority improves team performance’: authority for whom, exactly, to decide the schedule, the spending, or the technical approach? A project that runs out of control may show that the claim was too broad, or it may show that the staff could decide but could not get the information, or learnt of one another’s decisions too late. Only when these differences have been traced can a hypothesis be framed that the next case can compare: in a team with the relevant skills, information, and timely feedback, does adding a particular decision right shorten the wait for approval?
When a confounding factor is proposed, it too should be open to checking. If the instrument is suspected, look for a way to calibrate or compare it; if the conditions are suspected of changing, consult the records from the time. Adding one uncheckable reason after every failure robs the theory of its chance to meet a counterexample.
When to change the frame
A local repair sometimes solves the problem and sometimes makes the explanation harder and harder to use. If the same kind of failure keeps recurring, if each time another exception has to be added, and if the new account can only explain the past and offers no checkable expectation, then it is worth reconsidering how the problem was described in the first place.
Another strong reason is that an alternative explanation already exists which handles the new anomaly and also accounts for the results the old method got right. The comparison should weigh more than the novelty of the new account: how many extra assumptions it needs, which difficulties it resolves, and whether it would throw away reliable results already in hand.
The physics of light went through a change of this kind.
Nineteenth-century physicists knew that light behaves as a wave. Water waves have water, and sound travels through air or some other medium; when light crosses space, what is it that waves? The ether was once an important hypothesis put forward to answer that question.
In 1887 Albert Michelson and Edward Morley used an interferometer to send light out and back along different directions and then compared the interference fringes where the beams rejoined. They were looking for the shift in the fringes that the earth’s motion relative to the ether ought to produce, and the effect they measured was far smaller than was then expected. This put the existing arrangement, which used the ether to account for the earth’s motion and the propagation of light, in difficulty.34
Researchers went on revising their assumptions and improving the experiments. The work of Lorentz and others developed important mathematical relations that later physics was able to use. The process involved trial and correction, and the alternative account did not appear complete after a single experiment.
Special relativity later restated the relations between time, space, and the speed of light, allowing the phenomena to be handled from different basic assumptions. Change sometimes concerns how the question is asked, and goes beyond correcting a single value; the new theory must still fit the reliable observations already made, and must submit to new checks.
The researchers of the time had no finished history of physics to consult. They had to judge whether the difficulty in front of them could still be handled by the old theory; the new mathematical relations might be used to amend the old explanation, or might be taken up by a theory yet to come. The results we now sort into an ‘old frame’ and a ‘new frame’ were tangled together in the same inquiry while they were taking shape.
What admitting an error may cost
In the middle of the nineteenth century, the two maternity clinics of the Vienna General Hospital had different rates of death among mothers. The gap between the clinic where physicians and medical students worked and the one staffed by midwives kept Ignaz Semmelweis on the trail.
In 1847 a colleague died after being wounded during dissection work. Semmelweis connected the lesions found in his body with the disease of the mothers and suspected that medical staff were carrying some contamination from the dissecting room and passing it on when they attended the women. He introduced washing with a chlorine solution, and the death rate afterwards fell. The explanation he used, ‘cadaverous particles’, was not the complete knowledge of microbes and infection we have today.35
The success of the washing made a possibility that was hard to accept impossible to ignore: the staff had believed they were caring for the mothers, and their daily routine might have been spreading a fatal contamination. Responding to this evidence meant re-examining how the work was done, and perhaps facing the harm that past actions had caused.
The cost of admitting an error of this kind goes beyond changing a rule. Professional identity, relations with colleagues, and one’s estimate of oneself may all be affected. These costs help us understand why correction is difficult; whether to accept an explanation should still be decided by the evidence.
Institutions can make admitting error more feasible. They can, for example, allow workers to file records of anomalies and to suspend a practice that is in doubt, and set out clearly who is to check and who is to answer to those harmed. This keeps responsibility in place, and it means people need not abandon a report because the cause has not yet been fully proved, or because they fear being humiliated for making it.
Updating quickly is not necessarily better
The size of a correction should match the evidence. A single anomaly may come from measurement error or a chance event; when several independent studies repeatedly reach the same result, and that result directly contradicts a core prediction, a larger adjustment is needed. Even when the matter cannot yet be settled, the doubt can be noted and the search continued for data that would tell the causes apart.
Before people receive the same new report, they often already carry different experience and different judgements. One person once found a sampling problem in a report of this kind; another has repeatedly made accurate predictions by relying on it. That history may lead them to give the new report different degrees of trust.
Whether either reaction is reasonable depends on how relevant the past experience is to this report. Does it come from the same research team? Does it use the same method? Has the problem found before been fixed? ‘I have been misled before’ or ‘I have always trusted it’ does not, on its own, account for how much trust this particular report deserves.
So when people disagree, they can begin by asking what each side believed before and why, and then compare which of those reasons the new report has changed. Someone may have been deeply sceptical, may have raised their trust after reading it, and still fall short of the other person’s certainty. Looking only at whether they agree in the end misses an update that actually happened.
Setting out one’s earlier judgement lets others see where the disagreement comes from; allowing new and reliable evidence to keep changing it stops past experience from becoming a permanent reason to refuse correction.
That admitting error deserves praise is easy to say. The hard part is that the person concerned still has to judge whether they were in fact wrong. Change too little and people may go on being harmed; change too quickly and a practice that was reliable may be thrown away. The later results have not yet appeared, yet both costs have already begun to fall on someone.
Comparing the evidence, recording the judgement, and bringing others into the inquiry can give a decision grounds. This work still does not guarantee that we will choose the right moment for correction every time. A reliable method must also leave room for this possibility: that a person listens seriously to criticism, checks the data as well as they can, and in the end still makes a judgement that will need changing again.
Part Three — The Finite Mind
9 — What It Takes to Understand a Person
In heavy rain, the woodcutter and the priest shelter beneath the ruined Rashomon gate, still talking over the testimony they have just heard in a murder case. A commoner comes in out of the rain as well and presses them to say what happened. Akira Kurosawa’s film Rashomon then lets him, and the audience with him, hear several versions that cannot be squared with one another.
A samurai travelling through the forest with his wife met the bandit Tajomaru. The samurai died and the wife was raped. The bandit admits the killing, yet casts himself as the winner of a fair duel. In his version it was the wife who demanded that the two men fight, and only then did he cross swords with the samurai.
There is no duel in the wife’s account. She says that after the bandit left, her husband looked at her with contempt. She went towards him holding the dagger, then lost consciousness, and when she came round the blade was in her husband’s chest. The dead man, speaking through a medium, offers a third account: his wife had asked the bandit to kill him, and in the end he took his own life with the dagger. Did the bandit kill the samurai, or did the samurai kill himself? The two causes of death cannot both be right.
With each telling the film acts the events out again for the audience. The people are the same few people, but what they do in the forest changes. The woodcutter later changes his story too: he did more than find the body, he saw the fight. In his version both men shrink back, and the exchange they finally stumble into is a sorry affair, nothing like the heroic duel of the bandit’s telling.
A baby’s cry comes from under the gate, and the three men find an abandoned child. The commoner takes the child’s clothing; the woodcutter rebukes him, only to be challenged in turn: where did the wife’s valuable dagger go? The commoner suspects that the woodcutter stole it, and that this is why he concealed what he saw. The accusation gives reason to question the woodcutter’s testimony as well.
The woodcutter nonetheless decides to take the child home and raise it. The priest at first refuses to hand the baby over, and lets go only when he hears what the woodcutter intends.36
When the priest gave him the child, what exactly did he believe?
Leave the unfamiliar word where it stands, for now
In 1947 Thomas Kuhn, then doing research in physics at Harvard, began reading early scientific texts for a course he had been asked to teach. Aristotle’s discussion of motion left him deeply puzzled. Looked at from the physics that came after Newton, many of the claims seemed impossible to sustain; yet this was a thinker so acute in other fields. Why should he appear so utterly different here?
Kuhn recalled later that the turn came when he came to understand Aristotle’s vocabulary afresh. The motion and change Aristotle discussed took in the movement of bodies, but also growth, alteration of qualities and similar questions. If one reads only with the later physicists’ usage in mind, in which motion means a change of position, part of the original meaning is shut out. Once he had recognised this difference, Kuhn was better able to see why each stretch of argument was arranged as it was.37
At first Kuhn was, in effect, marking a piece of physics homework full of wrong answers; later the question itself changed. What had been called motion turned out to include growth and change of quality, which he had not put into the question at all. The progress in reading came when the concepts used to judge the text also began to change.
The reasons Haidt heard in India
In 1993 the social psychologist Jonathan Haidt went to Bhubaneswar, in eastern India, to do fieldwork. He was used to understanding moral questions in terms of harm, rights and fairness, and when he met certain local demands concerning purity, hierarchy and duty, he found it hard at first to see what they meant in the life around him.
Looking back, Haidt wrote that the daily life of the people there gradually brought him to understand how some of these rules connected with ideas of obligation, care and the sacred. He stopped seeing only restrictions on individual choice, and began to see what those who kept the rules believed they were upholding.38
The hospitality a visitor receives may make him look at community afresh, but it cannot speak for the lives of everyone the rules bind. Those who bear the restrictions have their own experience of them. The more one understands how a rule sustains relationships, the better placed one sometimes is to point out which people within it have always been asked to give more.
Now back to the child under the gate. The priest still does not know everything that happened in the forest, but the woodcutter’s readiness to raise the child gives him a reason to hand over its care. The question of the dagger has not gone away, and in front of him is a life that cannot wait indefinitely. Living with other people very often goes on inside this kind of understanding, one that never adds up to a single overall verdict.
Understanding another person also changes our relationship with him. Decide at the outset that someone is making excuses, and every clarification he offers afterwards may turn into further excuses; he senses the distrust and says less. The attitude the one who understands has taken up is already part of the behaviour he is trying to explain.
Only by giving the other person the chance to correct us can we learn from the answer something we did not know before. His account can still be checked against records of what he did and against other people’s experience. Mutual understanding takes shape in these exchanges; neither side can complete it alone by guessing more thoroughly.
After you know his answer
Knowing what a person supports is where understanding starts. Suppose, for example, that a team is discussing whether to trial a new shift rota, and someone is willing to back it because it can be tried for a while and reversed if it proves unsuitable. You know his position, and you also know one condition he leans on: whether the consequences can be undone.
You can then ask: if publishing the new rota would cost some people the care arrangements they now rely on, arrangements not easily restored, would he still support it in the same way? The question tests how much weight “it can be reversed” carries in his judgement, and it helps him notice a cost he may have overlooked.
Once this consideration is understood, it can be watched for in other decisions: where consequences are easy to repair, is there a case for trying first; where they are hard to undo, is more preparation needed? But another person may oppose the rota for quite different reasons, and refuse it even though it can be reversed. To understand him, one has to ask again what his reasons really are.
Is anyone aboard the boat that strikes you
Zhuangzi, in “The Mountain Tree”, shrinks the scene to a single collision. An empty boat drifts into a man’s boat and he does not lose his temper; if he sees someone aboard, he shouts at him to steer clear. He shouts and gets no answer, shouts more urgently, and in the end abuse follows.39
The boat still strikes, but the anger now has an object: a person who ought to have answered and did not. The event has been placed within intention and responsibility, and the feeling changes with it.
The parable goes on to speak of the discipline of emptying oneself to wander the world. A reader can first notice something in the collision itself: whether we grow angry often depends on how we read the other person’s intention. Did that man see the boat was about to hit? Was he able to steer clear? Until such things have been established, deliberate refusal to give way is only one possible explanation.
Establishing the cause can change the judgement of responsibility. Some people truly cannot control the boat; others know the risk and fail to take an action that was within their power. The damage from the collision still has to be dealt with, but who should answer for it depends on what could be known at the time, what could be done, and what was actually done.
The causes of an action, the reasons the person gives for it, and the reasons that would justify it to those affected should also be kept apart. Stress can help explain why someone lost control, but it has not yet answered how much those in the way should have to put up with. An institution that makes staff afraid to report errors accounts for part of their silence; how responsibility is shared still has to take in authority, risk and the choices that were available.
Public discussion in particular needs to beware of another short cut: treating the calm, fluent account as the more reliable one. A clear account is easier to check, but some people lack the vocabulary, some are frightened at that moment, and some have suffered precisely the kind of harm that makes the composure demanded of them impossible.
If someone says angrily “it happens every time”, one can check how many times it has happened while dealing with the one occasion already confirmed. Even if “every time” is inaccurate, the harm that did occur does not go away. Helping him set out what happened improves the quality of the record, and it stops the argument from stalling on an exaggerated word.
Knowing what a person fears can be used to comfort him, and it can be used to threaten him. An accurate understanding of his weakness does not come with a reason to exploit it.
Someone who does not want to go on talking now may simply be exhausted, or afraid of being punished for voicing an objection. Stopping the harm already confirmed, keeping the record, and letting the person rest and speak in safety are often more urgent than further questioning. He does not have to explain himself to our satisfaction before he qualifies for care.
The same word may be making different demands
Kuhn’s example concerns the meaning of a word. Everyday discussion runs into a further difficulty: even when everyone uses the word “fairness”, they may prize different arrangements. Some think whoever arrived first should be served first; others think whoever is in most urgent need should come first. Translating the word into another language has not yet made clear whom they are asking to do what.
Take the redesign of a park. “We see it differently” may contain any of the following questions:
| Where the disagreement lies | What needs clarifying |
|---|---|
| Data | Are both sides using usage records for the same period and the same area? |
| Definition | Does ease of passage mean walking, wheelchairs, or the speed of vehicles? |
| Causation | Which design would cause crowding, and what evidence supports that? |
| Scale | Have short-term disruption from construction and long-term consequences of use been kept separate? |
| Purpose | Is it transport, recreation, or some other need that is being improved? |
| Values and risk | Which costs are hard for which people to accept, and why? |
| Authority | Who is entitled to set the shared goal, and how do those affected take part? |
A single quarrel may be stuck at several of these points at once. More investigation can settle some disagreements over data, but it will not decide for the park’s users which kind of life is more worth having; and saying that values differ does not exempt a mistaken usage record from correction.
Put the table back against a concrete proposal. Take a hypothetical case: the park’s managers want to add night-time lighting to make it safer to pass through. One user doubts that the lamps, where they are placed, would reach the stretch of path where people most often fall. Another believes the lighting would work, but worries about the effect on the creatures that roost there at night. A third demands that nearby residents be involved in the decision first. All three may say “I object”, but the reasons that need answering are different.
The first doubt can be checked against where accidents happen and how well the lighting works. The second calls for more ecological and usage data, and it also involves how the various costs are to be treated. The third asks about entitlement and procedure in the decision; however well the first two studies are done, they have not yet answered who may decide on whose behalf.
Figure 9.1 One objection may fall on different relationships. This is the hypothetical proposal carried over from the table above, not an actual park survey. Drawn for this book.
State first whatever part can be agreed
Discussion does not necessarily reach full consensus. In the park lighting example, the parties may agree on the existing accident data, and agree that more lighting would improve one particular path, while still judging differently whether the ecological cost is acceptable. Separating what has been confirmed from what remains disputed saves the next discussion from starting over with the same batch of data.
Some shared criteria can also reach across professions. A product team wants an early launch and a safety team worries about accidents, yet both may agree that a small error which can be corrected quickly and an error which might cause irreversible harm call for different kinds of review. That agreement can become the reason for arranging a procedure, after which the two sides go on to discuss which risks belong in which class.
Shared criteria must themselves be open to checking. The two sides may still differ over what counts as “reversible”, or discover that the losses borne by one group have not been counted. The differences should then be set out explicitly, rather than everyone continuing to use the same word and pretending agreement has been reached.
If rankings of value or basic beliefs cannot for now be reconciled, the disagreement itself can at least be stated clearly. One person holds that a certain loss cannot be traded away; another is willing to compensate for it with other goods. This is a different problem from misreading the data. Only by admitting that no shared criterion yet exists can the parties look for negotiation, procedure, or some other acceptable way of living alongside one another.
The cost of changing one’s mind
An idea one read about only yesterday may be easy to change; a judgement bound up with professional reputation, group belonging or years of investment may be hard to let go. When persuasion meets resistance, besides checking whether the evidence is sufficient, one can also try to understand what the change would mean for the person concerned.
He may be worried about losing the support of his peers, or afraid that admitting one mistake will wipe out everything he has worked for. Sometimes, clearly separating “this judgement needs correcting” from “this person has no competence at all” is enough to let the discussion go on. Which actions caused damage and must be answered for should still be dealt with specifically.
A distinction is also needed between a reasonable decision made on the information available at the time and a lapse in which evidence that was available then was ignored. Both may need to change when new data appears, but what has to be acknowledged and improved is different in each case.
After changing his judgement, the person still has to go back to his own community and face those who trusted him, or were hurt, because of what he used to maintain. Correcting a statement may be only the beginning of that work.
10 — Borrowing Other People’s Thinking
LeetCode is a platform for practising programming. A problem may ask for the shortest route, or for a way to find one item quickly in a large body of data. Anyone who starts practising soon meets “data structures” and “algorithms”. The first studies how data should be arranged so that it is convenient to use; the second sets out the steps for getting a job done, and the reasons behind them.
The statement of a problem is short; finding a good method may have cost researchers many years. The layer-by-layer search for a path in Chapter 4 lets us lean on an existing proof to know when the search can stop, and why no shorter route has been missed. Someone who has learnt the method spends their time recognising the conditions of the problem. Someone who has never met it may be worn out simply from trying a few routes.
Trying for yourself has value. Where you get stuck, you can see what makes the problem hard, and you learn to recognise what exactly a method saves. But if never looking at existing results is taken as proof of independent thought, learning becomes a test of reinventing things. What humanity has accumulated is given no chance to work, while the individual’s stamina is asked to stretch without limit.
Faced with the hard questions of life, though, we easily forget this way of learning. Should a person keep every promise? What kind of exchange is fair? If a choice benefits me, is that sufficient reason to make it? These puzzles may have landed on us only today, yet philosophers have been arguing about them for a very long time. The distinctions, counterexamples and arguments they have offered will not live our lives for us, but they can help us see where exactly we are stuck.
A programming problem usually comes with its inputs, its constraints and its test for acceptance already given. In a discussion of fairness, even the test for acceptance may be in dispute. That makes borrowing ideas a matter for more judgement. Reading an argument, you may come away with a method; you may also discover only then that the answer you had been pursuing had ruled out certain people’s situations in advance.
Darwin reads a book about population
In his autobiography Charles Darwin recalled that selection in the state of nature had once puzzled him. He had already seen the power of artificial breeding, but he still needed to understand how, with no breeder doing the choosing, differences among living things could be kept.
In 1838 he read Malthus’s book on population. Population can grow faster than the means of living, and that problem made him think again about the struggle for existence among animals and plants. In an environment of limited resources, would certain differences make an organism more likely to survive and reproduce, so that those differences were passed to descendants more often? The differences he had gathered over long years of observation now had an explanation that could be worked on further. He still spent many years after that arranging evidence and revising his ideas before he published On the Origin of Species.40
Darwin read with a puzzle left over from long observation, and only because of that did the relation between population and resources turn, before his eyes, into a question in biology. Another person reading the same book need not have arrived there. Making the idea stand up afterwards meant going back to the evidence on species, variation and reproduction. The borrowing happened in this movement back and forth. As for how people should treat the weak, or how a struggle for existence should be conducted, none of this decided that for us.
When we read works of philosophy, we too can choose with a question in hand. You may accept a distinction an author draws about “freedom” while disagreeing with the policy he builds on it. To explain which step you have adopted, and why, you do not first have to become a disciple of some school.
But how does a book hand this ability to its readers? Darwin could say that his reading gave him inspiration. For those who come after, what is more useful is to see how he applied the relation between population and resources to variation among living things, and then to keep checking along that same step.
Leaving the reasons for the next person
Start with a very small example. Why is the sum of two even numbers still even?
You can try four plus six, eight plus twelve, and a few more pairs. To hand the discovery to someone else, though, there is another way of writing it. An even number can be split into two equal whole-number parts. Halve each of two even numbers and put one half from each together; the sum can still be divided into two equal whole-number parts. In symbols, write the two even numbers as 2a and 2b; adding them gives 2(a + b), where a and b are both integers.
The proof about even numbers does not record every detour the discoverer took. It leaves out the mood of the moment and the order of the trial sums, and keeps the relation that is enough for a stranger to redo it. This omission has a power of its own. A reader can point out that “any integer” cannot be substituted for “even number”, without first having to judge whether the author, as a whole person, deserves trust.
This is one point on which I set particular store by stated reasons. They make it possible for different people to meet at the same place and to disagree over the same definition, the same step of inference; corrections can therefore pass into the hands of people the original author never knew. When one person works something out once, the benefit need not stop with that person.
To achieve this, the discoverer has to spend more effort, not less. Writing down what you grasp at a glance is often more trouble than simply going on using it. The ease the reader receives contains the work an earlier person did for the sake of passing it on.
Some abilities are handed down mainly through demonstration and practice. In learning to knead dough, knowing that it should be “kneaded to the right degree” is not enough; the learner needs to touch different doughs, try the movements, and then ask someone skilled to point out the differences. Learning to recognise an abnormal noise in a machine likewise takes repeated listening to normal and abnormal examples, checking one’s judgement each time. Words can remind us to attend to the spring of the dough or to a particular stretch of sound; the actual feel, the discrimination by ear and the movements still have to be practised.
Sometimes the words stop right there, at “springs back to about this degree”. After that, teacher and student have to touch the same piece of dough together. The inference that can be written, the difference that can be pointed to and the movement practised by hand may only together make up the ability the newcomer actually learns.
Figure 10.1 How other people learnt, and then handed it on. Stated reasons and shared practice can interleave; checking and practice can uncover errors, but what is handed down may also go uncorrected indefinitely. Drawn for this book.
Mathematicians have carried this work of passing on further still, to the point where a machine can take part in the checking.
There is a kind of work in mathematics that sets out its reasons with unusual care. Lean is a tool for expressing definitions, propositions and proofs. People write what is to be proved in a precise form and then supply a proof; the system checks, by its logical rules, whether that proof supports the proposition as written.41
Where a proof on paper would say “from which it follows”, a reader can perhaps fill the gap alone. Handed to a tool of this kind, the omitted steps must be capable of being filled in as a checkable proof. The system can help complete some of the details, but whether the definitions are well chosen, and how existing results are joined up, may still take a great deal of work.
The even-number example above is short, and formalising it seems more trouble than simply seeing it. Once proofs grow long and the theorems they cite multiply, this checking becomes more helpful. A later worker can confirm which proposition they are citing and which premises it needs, then use it in a new proof. Nobody has to guess afresh, at every citation, which step the original author left out.
What Lean checks is whether, under the logic and axioms adopted, the proof given derives the proposition written down. That result still depends on the checking tool working correctly; whether the proposition is fit to answer a real question needs a separate check. If a model for studying a bridge left out an important loading condition, a proof inside the model, even one that passes the check, says nothing about how the real bridge would behave under that condition. The question from Chapter 5 sits right here: what was left out when the model was built, and does a later use happen to need exactly that?
Working a problem through once, properly
To ask whether a job is meaningful may be to ask about income, contribution, autonomy, the growth of a craft or a sense of identity; looking up satisfaction ratings for various job titles may not reach the thing you care about. The hard questions of life rarely run, as the even numbers do, from a clear definition straight through to the end of a proof. Borrowing the thinking of those before us can still begin by redoing a stretch of reasoning. Here is a hypothetical problem. You have promised to help a friend finish a piece of work, and circumstances have since changed. You hold that “having promised, you absolutely must not withdraw”, but carrying on would impose a clear burden. How should this principle be understood?
“Having promised, you must not withdraw” may express several different demands: no withdrawal under any circumstances; promises should normally be kept, but a major upheaval allows an exception; or the arrangement may be changed, provided you deal with the preparations the other person made on the strength of your promise. The original sentence does not separate these meanings, so the first question is which one you actually agree with.
Put the strongest version in a hard place first. If what was promised would itself harm someone, is there still a duty to perform it? If a major change arises that could not reasonably have been foreseen at the time, can the responsibility not shift at all? These counterexamples force the principle to narrow, while the weight of everyday promises still has to be explained separately.
Why a promise carries weight can be seen in the other person’s arrangements. Because you agreed, he may have turned down other help, set time aside or entrusted important work to you. Withdrawing now would cost him the options he once had, perhaps leaving no time to find a substitute. These consequences supply reasons to perform, to give early notice or to help repair the damage.
In this way, permission to withdraw and continuing responsibility can both hold at once. A major change may be enough to alter the original arrangement, but early notice, an honest account and whatever remedy lies within your power may still be required. After the principle and the case have been checked against each other, what emerges is a finer judgement: the promise keeps supplying reasons, and the weight of those reasons depends on what was promised, on what has changed and on the reliance that has already formed.
Only when conceptual analysis has come this far does the practical enquiry have a direction. What exactly was promised at the time? What arrangements has the other person already made? How heavy is the new burden? Could a different form of help be offered? These facts have to be obtained from the real relationship; thinking the principle through a little further cannot answer them in its place.
Where there were once only two options, grinding on or breaking faith, there are now arrangements that can be discussed with the friend: narrowing the scope of the help, moving the deadline, finding a replacement, or bearing a reasonable remedy after withdrawing. The friend may accept, or may point out that none of these makes good the loss. The analysis gives the conversation a clearer starting point; it has not agreed anything on the friend’s behalf.
The analysis above drew on conceptual clarification, counterexamples, the pressing of reasons, and the checking of a general principle against a particular case and back again. All of these are philosophical work that can be read about further, practised and criticised.
From there you can go deeper, choosing works according to the difficulty you have met. Peirce, in “The Fixation of Belief”, discusses how people settle their beliefs, and asks which ways of enquiring allow experience to revise what we originally thought. Dewey’s How We Think starts from concrete perplexities and discusses how people propose possible explanations and then check them.4243 The analysis of the promise above can go on borrowing tools from studies of conceptual clarification, of counterexamples, and of checking general principles and particular cases against each other.
What reading yields can be quite concrete. Two meanings that were once run together can now be kept apart; where there was once a single explanation, you now know there are rival explanations that can be checked. Each work makes broader claims of its own, and learning one of its methods is only a beginning. Redoing one of the author’s inferences with an example of your own, then changing the conditions to see whether it still holds, will help more than remembering the name of a school.
The practice should not end the moment a satisfying answer arrives. Change the situation and see what the method just used can do. Conceptual clarification can let people state their disagreement; it may not make them value the same thing. Weighing consequences can supply important information; it does not automatically turn rights and promises into figures that cancel one another out.
“Every voluntary exchange is fair” can be tried the same way. If one party’s alternatives have been deliberately destroyed by the other, is the consent that remains still enough to make the exchange fair? Pressing this one step shows which situations “voluntary” had been leaving out all along. Whether you are willing to revise the principle in those situations is what gives the counterexample its force.
Still, learning to redo this stretch of analysis is not yet the same as being able to read every argument that has come down to us. The promise just now was a hypothetical we had taken care to spell out; with an old piece of writing, what the author was answering at the time may not be written on that page at all.
Understanding a sentence means finding its question again
An instruction says only “handle this as soon as possible”. Every word is familiar, yet the task has not necessarily been made clear. Does it mean a reply before the end of the day, or dropping whatever is in hand at once? Does handling it mean acknowledging receipt, proposing a plan, or finishing the whole job? Are there checks that must not be skipped for the sake of speed?
The person who wrote it may have assumed that, since the two of them had just come out of a meeting, none of that background needed saying. Whoever takes over later has only those few words, and has to recover the question and the constraints that go with them. Understanding a text sometimes needs exactly this work: setting your own everyday usage aside for a moment and asking what the other person was responding to at the time.
The reader here needs a capacity for understanding other people: to find out what the other person knew at the time, what choices were open to them and which question they were answering, and then to try to reconstruct how they reached their conclusion. This is one layer of what “empathy” can include. Understanding an argument does not necessarily require sharing the author’s emotions, but it does require not treating outcomes you know about afterwards as things the author already knew.
In their study of communication, Clark and Brennan stress that common ground has to be built and updated within the interaction. Partners in a conversation gather evidence of how far they understand each other through responses, acknowledgements and repairs.44 Written material can also supply examples, definitions and context in advance; when the author is absent, the gaps may have to be filled by the reader’s other reading, hands-on work and discussion.
How much background needs filling in depends on how the material was written. A carefully laid-out textbook may be easier to understand than a hurried spoken explanation; the advantage of talking face to face is that you can ask at once. Whether communication has been adequate is judged by whether the receiver can find the necessary information and confirm that they have understood correctly.
Figure 10.2 How a reader finds out what the author meant. A reading is first proposed from the text and background material, then checked by asking, comparing or trying things out; new material found later may force a reading that once seemed smooth to be rewritten. Drawn for this book.
Reconstruction also has a trap that is easily overlooked. The smoother the explanation, the more likely we are to forget that it is still a guess.
In a series of studies, Eyal, Steffel and Epley compared imagining another person’s perspective with actually obtaining information from that person. In the judgement tasks the studies set up, asking participants to put themselves in the other’s shoes did not consistently improve the accuracy of their judgements about what others thought and felt; the condition in which perspectives were obtained through conversation did improve understanding.45 Imagining a few more details and knowing more about the other person’s actual situation do not have the same effect.
Back to “handle this as soon as possible”. We can propose several readings first, then look at the minutes of the meeting or simply confirm the deadline. With an old book, what can be checked is the surrounding text, the same author’s usage elsewhere, the materials of the period and other scholarship. Different objects call for different checks; only by giving your reconstruction a chance to be corrected do you avoid merely writing another handsome story on the other person’s behalf.
Stretch the distance to several hundred or several thousand years, and restoring the background takes far more work. Being able to read the symbols is a very small part of it.
The Yijing, the Book of Changes, is a body of texts that has passed through long handing down and interpretation. Each of the sixty-four hexagrams is made of six lines, and each line can take one of two forms, yin or yang; the hexagram statement is a written account of the hexagram as a whole, while line statements are attached beneath each line. Later interpretation gave further meanings to the hexagram images, the positions of the lines and the words. Knowing how six lines combine is only the starting point for reading the figures; the reader also needs to know how a particular interpreter used them to judge a situation.
The core hexagram and line statements, and the writings later called the Yizhuan, the Commentaries on the Changes, or the Ten Wings, took shape in different periods and passed through long use in divination and long interpretation. The dating of the texts and the manner of their formation are still under scholarly discussion.46 In reading, therefore, one needs to be clear which layer of text one is reading, and from which period, and which interpreter, a claim comes.
On a first reading, you can learn one interpretation that has textual grounding, then compare it with other readings. On what kinds of question is the same hexagram cited? Which hexagram or line statement do the reasons come from? What assumptions has the interpreter added? Doing this lets you check whether you have understood. Whether divination has any predictive power is a separate matter; it needs evidence that can tell successful predictions from failed ones, and familiarity with the text cannot stand in as proof.
Sometimes a reader who already has a mature method of their own meets an old set of figures and finds a new correspondence. Leibniz was a reader of that kind.
After developing binary arithmetic, Leibniz corresponded with Joachim Bouvet, a Jesuit in Beijing. The hexagram diagrams Bouvet sent him showed him a very attractive correspondence: match the two kinds of line to 0 and 1, and under a suitable arrangement and way of reading, the six-line combinations can correspond to six-digit binary numbers. Leibniz’s paper on binary of 1703 discussed these Chinese figures.47
This mathematical correspondence can be checked directly: two choices at each position, six positions, sixty-four combinations in all. Leibniz had found a way of understanding the hexagram diagrams through his own mathematics. The arrangement he saw, however, belonged to the tradition of Shao Yong in the Song dynasty; how ancient users understood those figures still has to be explained from the texts of their own time and other historical sources.
Two different results appear here. A new correspondence may help mathematical thinking; to claim that the arithmetical knowledge of the ancients has been recovered, historical evidence must be produced. Success in the first does not automatically accomplish the second. Even once the rules of the symbols are clear, there remains a great deal to learn about how their users understood them in ritual, in community and in daily life.
Someone reading the Yijing today may also want to borrow it for a problem in front of them. That requires setting out one’s own usage, so that others know what we are actually proposing.
In borrowing yin and yang, this book adopts a contemporary usage of very small scope: when faced with a one-sided judgement, first look for the complementary conditions, costs, counterweights and changes it has overlooked. This is an exercise in asking questions, and no generalisation about every ancient use of yin and yang.
Suppose a project has gone a long time without results, and “we should continue” and “we should give up” seem to leave room for only one answer. One can first find out what abilities and data have actually accumulated over that time, then reckon which options the investment has crowded out and which resources are near their limit. With that information, it becomes possible to compare whether to try once more, to change approach, or to end the project.
If, after swapping in a new set of words, we still know only that we want to continue, and have found no new reason and no checkable difference, the borrowing has not helped. Whether a method is useful depends on what more it made us see and check; the ancient name cannot vouch for it by itself.
The past a file name carries
“CON” is only three letters, yet under the file-naming rules commonly used in Windows it is no ordinary name. Along with PRN, NUL and others, it is reserved for devices; simply adding an extension does not necessarily turn it back into an ordinary file.
Raymond Chen, an engineer at Microsoft, has traced what such names were used for in the DOS era. Tools could handle devices in much the way they handled files, and the special status of the names had to survive conventions such as programs adding an extension automatically. By the time the system acquired fuller support for paths and directories, the expectations early programs had about names still had to be respected.48
To know which operation is restricted today, one still has to look up the rules for the interface, namespace and version in use. The history of how it took shape explains which dependencies to check; only the current documentation and actual testing tell us how the operation in front of us will run.
The use behind the name is on record and can be checked. Other origin stories that sound reasonable do not necessarily have the same footing. The keyboard we use every day is one example.
Why does the first row of the keyboard begin QWERTY? The common answer is that early typewriters, to stop the type bars jamming, deliberately made people type a little slower. The story is easy to remember, and it makes it easy to believe that today’s familiar layout came from a constraint that has since disappeared.
Koichi Yasuoka and Motoko Yasuoka, who have studied the early sources, have cast doubt on this popular version. Tracing the changes in the keyboard layout and the demands made by its early users, they argue that the problems telegraph operators met in transcribing Morse code provide an important clue to understanding how the layout changed.49 This is an explanation argued from historical sources, and it has not settled every detail of how the layout formed.
The keyboard story can therefore also be used to check how we accept explanations. “To avoid jamming, so they slowed typing deliberately” joins purpose, method and result very smoothly; to confirm that this really was the design reason at the time, one still has to find evidence of decisions and modifications. Having found another, more attractive version, one owes it the same check.
Once the sources are in hand, one still has to trace how that choice was carried forward afterwards. An arrangement that saved effort at the time may run into new difficulties only after a great many people have come to depend on it.
In some early computer systems, only the last two digits of the year were stored, so 1999 was written 99. Two characters fewer per record had real value where storage was expensive and data piled up in volume. The cost was left to later users: when 2000 arrived, should 00 be read as 1900 or as 2000?
If a system used the year to order events, or to calculate ages, interest or terms, the difference of a century amounted to far more than two missing characters on a screen. How the data were encoded, how the programs computed, and what assumptions the systems exchanging data with one another had made all needed to be checked.
In the late 1990s, governments and businesses invested heavily in inventories, patches, testing and contingency preparations. After the turn of the year, the large-scale disasters that had been anticipated did not generally occur.50 Assessing whether the preparations were worth it requires finding out which faults were actually discovered, which fixes removed risk, and whether the spending was proportionate. Looking only at the fact that the new year passed largely without incident tells us neither what would have happened without the fixes nor which measures had what effect.
What the two-digit year left behind was a choice that was once understandable, and a dependency that lived longer than expected. Tracing the history of how it took shape lets us know which places to check; it does not require us to keep things as they were for ever because there was once a reason.
A rule that is hard to understand may once have solved a problem that no longer exists. That is a hypothesis worth checking; it is not yet the rule’s history.
If we ask only “what reason would lead a reasonable person to do this”, it is easy to invent a plausible origin for whatever exists. The actual cause may equally have been an oversight, an imbalance of power, an accident carried forward, or several mutually incompatible modifications. To tell these possibilities apart, historical research needs the material that has survived; it cannot rely on how smoothly an explanation sounds.
An old comment in the code may preserve the original reason, or it may be out of date. How the program works today still has to be confirmed from current documentation and tests. Even when the original reason has gone, later programs may have formed new dependencies, and the consequences of removing a restriction have to be traced separately.
A rule may persist because its function is still there, or because of the cost of replacing it, vested interests, or the fact that nobody ever had the authority to change it. Tracing the past can separate these possibilities more clearly; deciding whether to keep the rule today still means weighing the evidence and costs in front of us.
A bridge lets us separate two kinds of question. Why the designers chose a particular safety factor, and why the code demands a particular test, are questions for the design goals, the known risks and the technical conditions of the time. How the bridge will deform or fail under a specific load is a question for mechanical analysis and the relevant testing. Understanding the design reasons helps to reveal which situations the model may have left out; the forces the structure actually bears do not change on that account.
Investigating the history of how something took shape and studying the properties of the thing itself can therefore help each other, yet they answer different questions. The first is particularly suited to recognising why programs, procedures, interfaces and classifications have the form they now have; to calculate a load-bearing capacity or find the limits of an algorithm, one still has to study the relevant properties and inferences. Sometimes the sources fall short, and the origin cannot be established within the deadline; then the unknown has to be written down clearly, and the decision supported instead by present-day tests.
Where to begin can be judged from the difficulty at hand. Asking “what is the essence of a file name” will not necessarily explain the special rule for CON; finding out what it was once used for can point to concrete compatibility problems. Turn to the load calculations for a bridge, and the history of the design cannot substitute for the data the calculation needs.
Learning from the aircraft story to check a different set of data
What would it mean to have learnt the aircraft case from Chapter 8? Remembering only “pay attention to what cannot be seen” may still leave you not knowing what to check when a new problem comes. More exact questions can be kept instead. What conditions decide whether a case enters the data? Would the cases that did not enter happen to change the answer we want?
Now apply the questions to feedback on a course. If everything collected is the evaluations of those who completed it, first find out who dropped out after enrolling, and why. Falling behind, running out of time and losing interest each mean something different for evaluating the course. Reading only the responses of those who finished, one cannot know the course’s effect on everyone who enrolled.
The two cases can borrow the same method of checking: recognise how the data were filtered, then ask whether the filtering bears on the question under study. Their specific causes differ, so a conclusion such as “reinforce the places with fewer bullet holes” cannot simply be carried across. How to obtain information from learners after they drop out still requires a survey designed afresh for the course.
This is the detail to keep when borrowing a method. Remembering only the aircraft, the bullet holes and the war, one easily assumes it is useful only for military matters; left with only “do not ignore the unknown”, one cannot guide a single enquiry. Explaining how the filtering process affects the data is what lets those who come later know which groups to compare.
If, when the question, the evidence and the conditions all change, you can still always use the same method to prove the answer you liked in the first place, it is worth going back to check: is there any result that would really make you change your view? If there is none, using the method may be nothing more than supplying reasons for a preference after the fact.
Understanding that can be carried over does not always come from agreement. Finding where an analogy fails may teach us a difference we had not noticed. Conversely, an example that holds completely, if it says only what we already knew, does not necessarily add to our capacity for judgement.
Ways of thinking accumulated like this can cross out of their original disciplines without our having to declare every field the same thing. Mathematical proof, historical tracing, shared practice and counterexamples each help us obtain something different. In choosing among them, we also gradually learn to explain why the case in front of us needs this one.
Those who come later will meet problems those before them never met. Learning a method is for having the ability to get that far; once there, sometimes the method itself has to change too.
11 — Thinking Has Deadlines Too
On 20 July 1969, Neil Armstrong and Buzz Aldrin were descending from lunar orbit towards the surface of the Moon in the lunar module, the Eagle. Mission control on the ground was receiving a continuous stream of flight data and talking with them by radio.
The guidance computer suddenly threw up a programme alarm. Armstrong reported the code, 1202, and shortly afterwards asked the ground again for a reading on it. Knowing which code it was did not yet let the astronauts decide whether to carry on descending. The module did not hang in the air while it waited for an answer.
What the ground had to answer at that moment was whether this alarm was interfering with any function the landing needed. A 1202 meant that the work the computer had scheduled exceeded the resources available, but the software had been given an order of priorities in advance, and after an overload it could still recover the guidance tasks that mattered. Whether the system could go on doing those tasks became the crux of the judgement. Why the extra load had appeared had not, at that point, been traced.
Jack Garman, who knew the software well, advised continuing. Steve Bales, responsible for the guidance system, assessed the alarm and the state of the vehicle and read it as safe to keep descending. Charlie Duke passed the answer up to the Eagle.
The alarm came back. Aldrin reported the same code, and said what data had been on the display when it appeared. The ground replied that they would be watching the difference between the two estimates of altitude. The descent went on, and so did the checking of status over the radio. A question that had just been answered had to be judged again as each new report came in.51
Only after the module had landed safely did the engineering teams have time to hunt down the extra load. Fred Martin, who took part, recalled that they examined the software, used the simulators, and checked the telemetry against the operating procedures, and traced the additional load to something connected with the rendezvous radar. Those findings bore on later flights; the fact that this landing had succeeded was no reason to stop asking.
The judgement that allowed the descent to continue was far narrower than “we understand the whole fault”. It rested on the kind of alarm, on how the software handled an overload, and on the essential guidance that was still running. The question the later investigation set out to answer mattered at the time as well, but nobody could demand that it be settled first before the module was allowed to fly on.
The same division of labour across time appears in ordinary life. A conversation needs continuous responses and cannot wait for every motive to be dissected; afterwards, even with a whole evening to hand, you cannot finish the foundational research on every concept before you permit yourself to form the next step.
Analysis paralysis is sometimes caused by something other than a shortfall in reasoning: every question worth pursuing has been promoted to a necessary precondition of the decision at hand.
Starting from first principles, taking assumptions apart and following them down to more basic conditions, comes into its own when an old framework no longer applies. But it cannot guarantee that you have reached a place with no presuppositions. You are still using concepts, rules of inference, observations, and some knowledge you have accepted for the time being.
The more immediate limit is cost. To decide how to practise a foreign language next week, you do not first have to settle the nature of language, the complete neural mechanism of learning, and the ultimate purpose of education. These questions deserve study; whether they belong to the necessary groundwork for this particular decision is a separate judgement.
If a method that has already been checked is to hand, and your situation broadly meets its conditions, using it first is usually more feasible than rebuilding every reason on your own. When the results turn out oddly, when the conditions are plainly different, or when the method’s premises run straight into your central doubt, that is the place to put more analysis.
Tracing things back to more basic conditions ought to help us identify which assumption has gone wrong. If each layer we peel away adds another requirement of the form “this must be fully thought through before we can begin”, the task of planning next week’s practice will never be finished. At that point we need to ask again: which of these doubts has an answer that would really change this plan?
Acting in the moment, checking afterwards, and learning over the long run make different demands on thinking.
Figure 11.1 Action in the moment, analysis afterwards, and long-term learning follow on from one another. The segments are not drawn to scale in time, nor do they imply that everyone should follow the same schedule. Drawn for this book.
High stakes alone do not settle that one should slow down either. Some high-stakes situations, precisely because time is so short, have to rely on trained, rapid responses. The questions that matter are what reliable abilities and materials were available at the time, what delay would cost, and whether the conditions can be improved in ordinary times.
Half the problem is in the environment
Herbert Simon studied how organisations make decisions and proposed the direction that came to be called “bounded rationality”. Real people cannot obtain all the information at once, list every option, and then compute the best answer at no cost. Decisions happen within limited time, knowledge and ability, and those conditions need to enter the explanation directly.52
One strategy he discussed goes by the name satisficing: you first set a requirement that would be acceptable, and once the search turns up an option that meets it you can stop, without exhausting every possibility. For instance, you might first establish that an arrangement has to be affordable and fit the time available, and then decide among the options that qualify. Set the requirement too low and you miss improvements that were worth fighting for; too high and you may never find a workable plan. The requirement can also be adjusted in the course of the search as new information arrives.
He likened rational behaviour to a pair of scissors: one blade is the structure of the task environment, the other the abilities of the agent. Looking at one blade alone, it is hard to explain the work that gets done. If the environment gives off stable signals, a simple rule may be enough; when the signals change, or when mistakes are costly, the same rule may no longer suit.
The same person, with clear records to hand, a usable method already learnt, and someone who knows the situation within reach, has many more ways forward than when facing vague data alone. The decision-maker is still the same person; the judgements they are able to make have already changed.
How a bounded analysis gets finished
Take a hypothetical case. You have to decide whether to change your current study arrangements next week. You read a great deal every evening, yet feel that very little of it is genuinely usable. Tonight you have only forty minutes to deal with the question.
What tonight needs is an arrangement you can try next week. As for why you have kept studying this way, that may involve habits many years old, and it will not all become clear in forty minutes.
The material may be too hard; you may be reading a lot and using little; and there is a further possibility, that you have in fact learnt something and simply never checked. These explanations lead to different arrangements, and it is worth trying to tell them apart first.
So you pick one concept you read recently, close the material, and try to explain it and to work a fresh example. If you cannot even state the basic meaning, the material or your prior knowledge deserves another look; if you can state it but cannot use it, the practice that follows has a clearer direction. It may also turn out, once you try, that you can use it better than you had supposed.
Carry the case further. Suppose the check shows that you can explain the concept but do not know how to use it when a problem alters the conditions. Next week you could provisionally keep half of your reading time and give the other half to working through examples with variations. This adjustment targets the gap just exposed. If the check had instead shown that the basic concept was unclear, the priority would have to shift to shoring up the foundations, and the same timetable could not simply be applied.
A week later, check again with a new task of similar difficulty. If you still cannot apply the concept, the difficulty of the examples, the feedback, and your prior knowledge all need looking at again; if there has been progress, then decide whether to keep this allocation. Which step to take first, and how to judge the next, are already written into the arrangement. Whether it works is left for the actual learning to answer.
The remaining doubts can still be written down: whether the long-term goal is clear, whether the choice of material is too scattered, whether your present job is breaking your study time into fragments. They do not lose their standing because they went unresolved this time, but neither do they all have to stand in the way of tonight’s decision.
If the next check is likely to change an important decision, or to avert an obvious loss, there is usually reason to go on checking. If the feasible choices in front of you are the same however a minor detail turns out, spending a great deal more time chasing it may not pay.
But “would it change the decision” is not the whole of it. Some information helps you carry out the same decision, showing you how to do it better; some inquiry, though it does not alter today’s choice, builds capacity for a question that will keep coming back. These kinds of value have to be counted too.
On the other side lie the time, money and attention the checking itself consumes, and the cost of delay. In pursuing a better decision, you cannot assume that the process of arriving at it is free.
Bringing the time and resources that thinking requires into the evaluation connects with research on bounded rationality and resource rationality. The resource-rational analysis proposed by Falk Lieder and Thomas Griffiths studies exactly this: the effectiveness and the cost of different cognitive strategies under limited computation.53 For our question it offers a useful line of inquiry. Beyond comparing which answer is better, it also compares whether the way of getting to an answer is worth it.
Once we begin to calculate “is this still worth thinking about”, we may go on to ask “is it worth calculating whether this is still worth thinking about”. If every level is required to have a complete guarantee from the level above it, analysis paralysis has merely moved house.
In practice this needs some starting points that are open to revision: adopt a time budget proportionate to the task, check first the information that will swing the main differences, and keep a limited number of chances to look back. When you already have reason to believe the allocation of time has gone wrong, adjust it then; there is no need to re-prove the whole philosophy of allocation before every piece of work.
If the analysis keeps circling the same set of reasons, that is the moment to stop and look something up, try it out, or ask someone; if a cheap and reliable way of checking the key doubt is still available, it may be worth spending a little more time. Whether to continue should be decided by what can be learnt next.
Please also keep some time for inquiry with no immediate output. Philosophy need not always be in the service of tomorrow’s to-do list. Only be clear at the outset which you are doing: solving a problem with a deadline, or allowing a question to change you slowly. Both activities can have value. It is confusing them that makes it easy to lose the freedom to act and the freedom to think at the same time.
A choice you cannot go back and redo
A study arrangement can be adjusted again next week; some choices cannot be fully withdrawn. After accepting a job, you can check whether the duties match what was described, but you will not also live through the life in which you did not accept it. The other road leaves behind no result waiting for you to go back and read.
If things go well afterwards, that does not prove the other choice would necessarily have been worse; if difficulties come, the outcome alone cannot show that the decision was rash. What can still be traced is which expectations fell through, which information was off, and whether a check worth doing was missed at the time. Asking about actual working hours before the deadline, talking it over with the person you share caring duties with, and finding out what the day-to-day responsibilities are come closer to those questions than turning a few job titles over and over in your mind.
These efforts give a choice something to rest on; they do not spare anyone the choosing. While the module went on descending, the ground was still receiving new data. Some answers only begin to become obtainable after events have already moved on.
12 — What We Know Before We Can Say It
A fire crew entered a house to deal with a fire that appeared to have started around the kitchen. The crew put water on it in the familiar way, but the fire did not respond as the commander expected.
When the decision researcher Gary Klein interviewed this commander, he heard an account that he kept returning to with further questions. The commander had suddenly felt that something was wrong and ordered everyone out. After the crew had left, the floor where they had been standing collapsed. The real fire was in the basement. They had been standing above it the whole time.54
At first the commander called the judgement a sixth sense. If the matter stopped at that name, nobody else would know whether it could be learned, or when to trust it next time. Only as the interview kept working backwards did several cues gradually surface: the heat in the room was out of proportion to the fire he could see, the sound was unusually quiet, and the water was doing less than expected.
In that moment, he had not first arranged these into a complete argument about the basement. A few discrepancies had already combined into the feeling that they needed to get out.
These cues were recovered step by step in the interview afterwards. A researcher can follow them up, but cannot take the full explanation given in the interview and put it back, unchanged, into the commander’s mind before the retreat.
The scene did not simply present one extra, conspicuous danger signal. Some responses that would normally have occurred failed to occur: the fire did not die down as expected once water was on it, and the sound and the heat did not match. The commander noticed these differences first, and explained them step by step only in the interview. Had he been required to name the differences, quantify them and complete the argument before they were allowed to affect his decision, the time to get out might have been lost.
We often put “fast, intuitive, emotional” on one side and “slow, rational” on the other. These words describe different things. “Fast” refers to how much time was taken. “Intuitive” usually means that the answer arrived first, without the person being aware of any step-by-step reasoning. Emotion includes feeling, an appraisal of the situation, and a disposition to act. “Rational” sometimes means explicit inference, and sometimes means that a judgement responds appropriately to reasons. To compare two judgements, we first need to know whether we are comparing their speed, the way they were formed, or the adequacy of their reasons.
A quick answer may come out of years of practice; an analysis that lasts for hours may spend the whole time finding excuses for an existing prejudice. Emotion can prompt a person to notice a harm that has been overlooked, while reasoning can examine the cause of the harm, who is responsible, and how to respond. To evaluate them, it is not enough to ask which arrived first or which looks calmer.
Intuitions also differ in how they are formed. Some skills are learned from explicit rules and, once practised, no longer need to be recited step by step; other powers of recognition are formed mainly through repeated exposure, imitation and feedback, and the learner never wrote out a complete set of rules in the first place. With the former, one can trace the steps that were once learned; with the latter, one has more need to compare experience against outcomes. The fact that both arrive quickly does not mean that beneath each there must lie an argument waiting to be recovered.
Information we cannot yet put into words
In 2007 a group of researchers ran an experiment in which odours that participants could not consciously detect were paired with ratings of how likeable neutral faces were, and observed that under particular conditions the odours affected the ratings. This supports a limited claim: some sensory influences can take part in a judgement while the person is unable to report clearly where they came from.55
It does not show that such influence is more accurate. The odour in the experiment was no evidence of whether the person in the photograph deserved to be liked; in this setting, an undetected influence may even pull the judgement away from the very thing it was meant to assess.
A feeling whose source cannot yet be explained can serve as a starting point for checking. Someone who hears that a machine sounds wrong, for instance, can first record the sound, compare it with the normal state, and then check whether it goes with a fault. That the person cannot say which frequency has changed does not make the difference unreal; that the feeling is strong does not remove the need to compare.
Explicit reasoning needs access to the relevant information before it has any chance of judging well. If the only input is “the machine is still running”, no amount of careful inference from that sentence will produce a stretch of abnormal sound on its own. Sensory observation, instrument readings and hands-on operation can supply details that never entered the analysis.
The reverse also holds: explicit comparison can find what feeling has missed. Only when Chapter 5 set out two groups of numbers with the same average did we see that one group contained values below twenty; the impression that “average performance is about the same” could not have answered that question. Feeling and inference can each miss things, and each can make new differences visible.
An operator hears the fault first, measurement then confirms what kind it is, and a few sessions of listening together lead a newcomer to start noticing parts of the sound she had not heard before. By this point, feeling and explicit comparison have altered each other, and the ability that results is hard to credit wholly to either side.
Figure 12.1 How feeling, inference and checking help one another. Actual outcomes, comparison of examples and independent measurement can all change the original judgement; where feedback is delayed or incomplete, one may never learn where the error lay. The figure shows learning relationships, not the layout of brain regions. Drawn for this book.
After checking of this kind, the operator may know more quickly which part to inspect the next time a similar sound appears. Where the measurement did not support the original judgement, that record needs keeping too, so that we do not remember only the times the guess was right. Learning over the long run can improve the response in the moment, without the whole of that learning having to be redone before each action.
The ball is still in the air and the feet are already running
A baseball is hit towards the outfield. The fielder looks up to track it and starts moving at the same time. If we say his job is first to obtain the ball’s initial speed, its spin and the wind, and then to work out where it will land, we slip easily into thinking that the body is just quietly completing a ballistic calculation there was no time to say aloud.
Research on catching has offered another kind of explanation. The player uses the continuously changing relationship of the ball within his field of view, adjusting as he runs, with no need to compute a fixed landing point first and then run to it with his eyes shut. The models researchers have proposed include keeping the ball’s visual trajectory linear in some respect and regulating a particular optical acceleration; different models come with different conditions and testable predictions.56
The player’s movement also changes the trajectory he will see next. Researchers therefore need to observe how visual information and movement pass back and forth before they can compare the models. The process of catching does not sit waiting behind some complete calculation performed before the action.
Walking with a nearly full glass of water makes the same adjusting-as-you-go easy to notice: you see the surface tilt, and hand and step correct accordingly. Studying activities like these means observing what information people pick up while acting and how they respond; looking only for a comprehensive plan finished before the action began may miss the important part of the process.
Even Cook Ding slows down
In “The Secret of Caring for Life”, in the Zhuangzi, Lord Wenhui watches Cook Ding cut up an ox and is astonished at his movements. Hand, shoulder, foot and knee work with the knife like a performance with its own rhythm.
Cook Ding says that when he began, what he saw was the whole ox; after a few years, what stood before him was no longer a single, complete object. He runs the knife along the gaps between sinew and bone, rather than forcing his way through. The story pushes this skill to a point that is almost beyond belief: the knife has been in use for nineteen years and has cut up several thousand oxen, yet its edge is still as if freshly ground.
Where sinew and bone knot together and the knife is hard to place, he still becomes wary. His attention gathers, his movements slow, and he moves the blade by the smallest degrees. Only when it is done does he withdraw the knife.39
Cook Ding’s skill lies in doing the ordinary work quickly, and equally in slowing down when he meets a difficult place. The nineteen-year blade carries the exaggeration of a parable, but the pause raises a question that can be studied: what kind of experience lets a person carry a familiar movement through smoothly and also notice early that this time calls for special care?
In 2009 Kahneman and Klein discussed expert intuition together and set out two important conditions: the environment must contain stable cues that can be learned, and the person judging must have had enough opportunity to learn them through experience and feedback. They also cautioned that a subjective sense of certainty is, in itself, no guarantee that a judgement is correct.57
If errors are never pointed out, twenty years may do no more than make a certain kind of guessing more fluent. When feedback comes too late, or when outcomes are obscured by luck or by standing, a person may not even be able to tell which judgement went wrong. Equal years of service can hide very different histories of learning.
Recognition exercises in which the answer is withheld, records of predictions set against outcomes, or a measurement of another kind can make this difference visible. Someone good at telling sounds apart will not necessarily be good at estimating quantities, and a cue that was reliable may fail in a new environment; “having intuition” has never been a qualification that covers all of a job.
After understanding, there is still some way to go
Understanding a claim when you hear it and being able to use it when needed are different outcomes of learning. You may agree that “one failure does not condemn the whole person” and yet want to defend yourself the instant you are criticised; you may know that past investment cannot be recovered and still find it hard, facing a project you have worked on for three years, to decide only on the costs and opportunities ahead.
This gap can point to what needs practising next. Being able to explain a reason in a quiet moment tells us only that we have understood it; to recall it in the relevant situation and act on it, we also need to recognise the moment, remember what to do, and sometimes learn to bear an uncomfortable feeling.
In his 1982 work on skill acquisition, John Anderson proposed a theory in which explicit knowledge gradually gives rise to procedural ability. When first learning certain cognitive skills, people need to recall the rules one after another; with practice, execution may become faster and no longer require every step to be rehearsed silently.58 The theory helps us study how certain skills become fluent; changes in emotion and belief still require an examination of the learning processes proper to them.
Reflection therefore has a long-term use: choosing which practices are worth keeping, arranging for their repeated use and checking, and making them easier to carry out next time. Someone who, in an argument, always hears a single objection as total rejection can reread the exchange afterwards and practise separating the specific criticism from a verdict on the whole person; next time, they can first confirm which point the other person objects to, and then answer. If the practice helps, what changes may be the difference noticed first, rather than merely one more reminder committed to memory.
Moral judgement faces a similar demand. Some harms that call for a timely response cannot wait to be seen until a full debate has run its course each time; we hope that a concern which has passed through reflection will gradually shape ordinary attention and reaction as well. But habit in itself carries no guarantee of right or wrong: a group’s ways of discriminating, or of shifting responsibility, can be learned to the same high degree of fluency.
Ask the expert when they would stop
When a learner asks a master “how did you know?”, the answer is sometimes just “you can tell at a glance”. The question can be put differently: which detail made you change what you were doing? If that detail had been different, would you have carried on? Under what circumstances would you ask someone else to help check?
Questioning of this kind is one of the main ways Klein studied expert decisions. Rather than asking the interviewee to produce a complete theory on the spot, it follows a single actual decision and picks out which cues would have affected the choice. In the firefighting case, the heat, the sound and the fire’s response to water were exactly what the interview went on to pursue.
Having learned these cues, the newcomer still has to compare and practise in the relevant situations. An interview can tell him where attention is worth directing; a spoken account alone cannot give him the same experience. If the account can be set beside records made at the time and the outcomes that followed, it also becomes easier to separate what the interviewee noticed then from the explanations that formed only afterwards.
Leave a record before the outcome arrives
Looking back, we easily mix the outcome we later learned into the feeling we had at the time. A study published in 1975 by Fischhoff and Beyth examined this using judgements made before and after Nixon’s visits to China and the Soviet Union. Before the visits, participants estimated the probabilities of a series of possible outcomes; afterwards, they recalled how they had estimated them. In recollection, what had happened tended to become more foreseeable than it had originally seemed, and what had not happened seemed to have been more doubtful all along.59
If a record from before the outcome can be kept, there is a chance of catching this rewriting. For important judgements where time allows, you can write down first what you noticed earliest, what you expect to happen, and how confident you are; where the reasons are still unclear, record that as it is. Then, when you check afterwards, reasons that occurred to you later will not all be counted as things you knew from the start.
A record preserves the original judgement, but it does not automatically supply every outcome. Take a hypothetical manager who notes “this applicant may not be suitable” and therefore does not hire the person; however complete the log, it has not produced that person’s performance after joining. To compare prediction with outcome, one still has to identify which kind of data is missing.
Some outcomes are not determined by our own choice and could always have been tracked. A share that was considered but not bought, for example, will still have a market price afterwards. If a specific price expectation and time horizon were recorded beforehand, they can be checked later, rather than remembering only what was actually bought.
Other comparisons require actually putting different arrangements in place. A/B testing of a product exposes different groups to different versions, so the effect of a particular change can be compared. This design has its costs; where people’s opportunities or treatment are involved, fairness, consent and tolerable risk limit how one may experiment. Wanting to improve one’s own predictive ability is not sufficient to justify handing the costs to other people at will.
For still other questions, what is missing is the outcome of the same person taking the other road through the same stretch of history. The job that was not accepted cannot afterwards become another life one has already lived. Other people’s experience under similar conditions, broader statistics and relevant research can narrow the uncertainty, but they will not restore the individual outcome that never occurred.
Before checking, then, first establish which question this data can answer and what it leaves unknown. How much is worth investing must be weighed together with the cost of obtaining the information, the consequences of a wrong judgement, and how many future occasions there will be to use what is learned. Each of these can add to the reasons for checking, but “it matters a great deal” or “it will come up again” is not by itself enough to conclude that any expensive investigation is worth it.
Emotion shapes what we take to be a problem
Why something becomes a problem for us often has to do with emotion. The indignation of seeing someone humiliated can make behaviour that had passed as a joke worth questioning; concern for a person can keep us attentive to difficulties he has not spoken of. Emotion affects answers, and sometimes it has already affected what we are willing to ask.
This practical role also needs to be identified case by case. Anger may notice an injustice, or it may mistake frustration for another person’s malice; shame may prompt reflection, or it may stem from a group’s demands that do not deserve acceptance. Emotion needs to be understood and checked, but the checking should not assume that the only acceptable result is for the emotion to disappear.
You can discover that your anger was attributed to the wrong cause and still keep your original sensitivity to a certain kind of harm; you can also, through someone else’s account, learn to respond to experiences you never used to care about. Changes like these do more than speed up the response in the moment. They alter what you will notice in future, what you will remember, and what you will be willing to check.
Care keeps a person observing a difficulty over time; liking sustains long practice; disappointment drives a person to re-examine what they had expected. Feelings such as these can take part in understanding and learning over long stretches, and their role goes beyond quickly offering a guess before the analysis begins. The judgements they guide still have to be checked, but they cannot be evaluated by calmness alone.
Within these long-term changes, reflection alters what is felt next time, and feeling can make the old reasons come to seem insufficient. After you have understood what someone else went through, a kind of joke that once seemed harmless may no longer raise a laugh. That change does not have to be maintained by silently rehearsing the whole argument every time.
The conditions of learning play their part too. Comparing similar examples, receiving timely feedback and adjusting practice beforehand make differences easier to recognise in the moment. Whether the work allows time for rest, whether the people around you are willing to point out mistakes, and which cues the tools display, in turn set limits on whether these changes can happen at all.
If the present moment is abnormal and time is short, the practised response may still be the most usable capacity to hand; when there is time, one can stop to compare and check. Both may be the products of this joint learning.
What we change is sometimes a sentence we believe, and sometimes what we notice at first glance. Once the second kind of change has taken place, a person may no longer remember that it once took so much time to learn.
13 — One Body, Many Kinds of Regulation
Close your eyes and you will most likely still know where your right hand is: whether the arm is bent, stretched out, or resting on the back of the chair. This sense of position and movement is seldom noticed on its own. It is what lets the limbs keep working while the eyes are busy looking somewhere else.
In 1992 the researchers Jonathan Cole and E. M. Sedgwick reported on an unusual participant. He could still move voluntarily, and he retained some sensation of pain and of heat and cold, but below the neck he had lost most of the sensory input concerned with light touch and the position of his limbs. Being able to move a muscle and being able to feel how one is moving turned out, in his case, to be clearly different things.
Asked to compare weights while watching his forearm move, he could still tell apart quite fine differences; with his eyes closed, the ability fell away markedly. Certain postures and simple repeated movements could be kept up to a limited extent, but new movements needed visual feedback.60
Movements that ordinarily need no particular attention required him to keep watching, adjusting and relearning. Vision made up part of the missing information about position, and in doing so it took up attention that could otherwise have been turned elsewhere.
Other kinds of work need even less in the way of step-by-step commands from the nerves. When the edge of a sheet of paper cuts the skin, tiny blood vessels are damaged and blood seeps out, and the local process of stopping the bleeding is already under way. You can notice the wound, deal with it or ask for help, but you do not first approve each protein reaction in your mind.
When a vessel is injured, local signals, platelets and clotting proteins take part in forming a clot. This is a set of physiological mechanisms. It does not have to pass through conscious judgement first, and it is not a neural reflex in which a message travels to the spinal cord before an order to stop the bleeding is sent back.61
Platelets are small cell fragments in the blood that take part in stopping bleeding; the reactions of the clotting proteins help to form a fibrous mesh. The local reaction is triggered and amplified, and it is also shaped by mechanisms that limit it and clear it away. To explain when a clot forms, and why it does not go on spreading without limit, one has to study how these chemical reactions act on one another.
Spinal reflexes, by contrast, do involve neural circuits. Some responses can be organised without waiting for a conscious decision, while still remaining open to modulation by other neural activity. That a response goes ahead without permission from present awareness does not mean it is forever cut off from the influence of the brain.62
The brainstem connects the cerebrum with the spinal cord and takes part in vital functions such as breathing, as well as in a great deal of signal processing; the cerebellum takes part in the coordination and adjustment of movement, among other functions; and different regions of the cerebrum share in sensation, memory, language and planning. These parts are extensively connected, and many activities have to be carried out across regions together.63 To explain a piece of behaviour, one usually also needs to know which information is passed on and how, and which activities modulate which others. A list of organ names is not enough.
Figure 13.1 sets the two kinds of process side by side. The left shows the local stopping of bleeding; the right shows the interplay between sensory information, neural circuits and responses. Both sides can operate without a conscious decision in the moment, yet the particular ways in which they are triggered, regulated and limited differ.
Figure 13.1 The functional relations between different mechanisms. The figure compares division of labour and mutual regulation; it is not a complete anatomical diagram, and the arrows do not correspond one by one to neural pathways. Clotting, reflexes and conscious activity keep their separate mechanisms. Drawn for this book from the literature cited in this chapter.
Coughing shows mutual regulation
In a quiet room, when the urge to cough rises in your throat, you can sometimes hold it back for a while and sometimes cannot; you can also cough deliberately to catch someone’s attention. From the outside all of these are coughs, yet the processes that set them off and shape them are not quite the same. Research therefore has to separate the stimulus, the urge to cough, the number of coughs actually produced, and the activity that goes on during deliberate suppression.
Functional brain imaging studies by Stuart Mazzone and colleagues compared coughing with the suppression of coughing, among other conditions, and observed different patterns of brain activity, which supports the view that the control of coughing in humans involves networks above the brainstem. Differences in activity seen in the images help in studying the processes concerned, but they do not amount to establishing, from the images alone, the complete causal function of each region.64
The attempt to hold a cough back does affect it, and there are also times when it cannot be held. Being able to take an active part in regulation has not turned the body into a procedure that waits for one’s approval every time.
People can also change themselves through external things
Deliberate adjustment does not happen only inside consciousness. You can change how you practise, arrange the environment you sleep in, lean on tools, and you can also alter certain physiological processes through medical intervention. These measures work in different ways and call for different bodies of knowledge; they cannot all be treated as another name for the will.
A small randomised trial by Alyn Morice and colleagues in 2007 shows what needs to be kept apart when an external intervention is assessed. The study treated chronic cough with morphine and observed improvement on some symptom scores; a citric acid cough challenge in the same study, however, did not show a significant change.65 The two methods of measurement produced different results: one recorded symptoms, the other observed the coughs elicited by an experimental stimulus.
A threshold is the level a condition has to reach before a given response begins to appear. To say that a drug has raised some cough threshold, one has to state what the stimulus was, how the response was measured, and how before and after the intervention were compared. Coughing less in daily life, or feeling more comfortable, is not by itself enough to prove that a stronger experimental stimulus is now needed to bring on a cough. Only by keeping these results apart can one know which kind of improvement the study actually supports.
A person can decide to accept an intervention without having to direct in person every physiological response that follows it. Researchers identify what a substance does and design trials, medical workers assess whether it suits a given case, and institutions affect whether a person can obtain help at all; the change that finally takes place in one person’s body has depended on a great deal of work that was never inside that body.
This gives “changing yourself by your own efforts” a second meaning. A learner can choose the setting in which to practise, ask others to correct them, and use tools that issue reminders; these arrangements in turn gradually change their habits and their judgement. It is the present self that makes the arrangements, yet what the self can later do, and what it readily notices, will be shaped by them. A person’s agency can extend through external conditions, and there is no need first to assume an inner commander in charge of the whole body.
Which changes count as regulation
At this point it is tempting to call every natural change “information processing”. But if a falling stone, clotting blood, catching a ball and debating a regulation are left with only one name between them, the very differences that were worth understanding disappear.
Start by comparing a stone with a thermostat. Both obey the laws of physics, but the thermostat has a sensor, a temperature setting and a switch: when the measured temperature departs from the setting, the device changes the heating. By adjusting the setting or disabling the sensor, one can check how each part affects the outcome. A stone falls, and has acquired no such set of measurement and response in doing so.
Nor is clotting simply another thermostat. It includes the triggering, amplification and limitation of a local reaction, while the control of movement draws on many kinds of continuously changing sensation. Learning may further change how one responds in future. To compare these processes, one should point out which differences they detect, how those differences alter activity, and how the outcome feeds into later responses.
Feedback can also arrive too late, or amplify the original deviation. A thermostat that keeps heating on the basis of an out-of-date temperature may fail to stop at the right moment; two modules that take turns undoing each other’s changes may leave a piece of work being altered back and forth. The name feedback carries no guarantee that it helps. One has to check what it takes in and when it changes activity.
Letting fixed procedures carry out familiar tasks, letting the parts that can adjust their strategy deal with new situations, and then having people handle certain exceptions, is a division of labour that can be studied. Which part should take over depends on the information it receives, the time available to it, and the capacities it has already demonstrated. If a person cannot see the original situation and has only a few seconds to respond, writing “a human is responsible” does not supply the missing capacity.
What a whole can do may be something no single part can do on its own. Studying such a capacity means looking at local activity and at how the parts affect one another at the same time; calling one part the brain or the command centre still leaves one having to return to these concrete relations.
Once we cross over into society, the participants can also object to the goal. Workers protest, and their disagreement carries reasons that need to be heard; it cannot be treated across the board as a deviation awaiting correction. The body’s division of labour can prompt questions, but it cannot provide legitimacy for anyone’s right to demand anyone else’s obedience.
Part Four — Understanding the World Together
14 — Letting Other People’s Findings Change Decisions
Some people go to a bank in order to pay their salary into an account. They ask for no further services and agree to no other products, yet new accounts appear in their names.
Why would a bank open accounts for customers who know nothing about them? Besides serving existing needs, staff were also required to sell more products to the same customer. The number of accounts is easy to count; it can be recorded as performance, set as a target, and tied to rewards. If every added service came from the customer’s own choice, the count might well show that business had grown.
In 2016 the U.S. Consumer Financial Protection Bureau took enforcement action against Wells Fargo. The regulatory documents record that employees, in order to meet sales targets and earn incentives, opened accounts without customers’ consent; some of these operations also moved funds from customers’ existing accounts into the new ones, leaving some customers to bear fees. The problems the regulator pointed to included sales incentives and inadequate oversight.66
The accounts really were created and the transactions really were recorded, but the thing those numbers were originally meant to stand for did not happen: a customer choosing a new service because they needed it. For the bank’s evaluation there may have been one more sale on the books; for the unwitting customer there may have been fees, and paperwork that had to be sorted out.
A customer’s consent is itself a necessary condition for the service to exist at all. Besides counting the accounts opened, the institution needs to let customers find out what has happened, stop services they never agreed to, and deal with the losses already caused. Raising the efficiency of account opening cannot take the place of these requirements.
What a tail actually proves
In 1902 Hanoi, under French colonial rule, faced a rat infestation and the risk of disease. The newly built sewers brought modern sanitation to parts of the city, and also gave the rats a space in which to move and breed with ease. The colonial authorities organised a rat-killing campaign and paid by the piece.
Counting needs evidence. Large numbers of rat carcasses are hard to transport and tally; a tail is far smaller, and seemed enough to show that one rat had been disposed of. So the tail became the proof that earned the bounty. The historian Michael Vann traced the campaign through the colonial archives, and what he found in its aftermath included live rats without tails, and activity that supplied rats for the sake of the reward. A tail cut off could be exchanged for money; the rat that remained had not necessarily died.67
The tail was originally meant to prove that a rat had been eliminated; once a tail could be exchanged for money on its own, the supply of tails could come apart from any reduction in the number of living rats. Vann’s research also sets these practices back within the residential divisions, labour and power of a colonial city. Who designed the payment scheme, who could earn an income from it, and who bore the infestation all shaped the campaign, and the question cannot be reduced to whether individual rat-catchers were greedy.
The bank’s accounts and Hanoi’s tails both show us a gap in time: before a metric becomes a target, it may be connected to the thing we care about; once the target is announced and people begin acting on it, the original relationship needs to be checked again.
Several researchers have studied changes of this kind from different fields. Charles Goodhart, discussing monetary management in the 1970s, noticed that a statistical relationship may change once it becomes a point of leverage for policy; Donald Campbell studied the distorting pressures on indicators used in social evaluation; and Marilyn Strathern, in a discussion of university assessment, set out the problem that arises when a measure becomes a target.68 The relationship observed when a metric was set will not necessarily persist unchanged once rewards and penalties have altered behaviour.
Figure 14.1 How evaluation changes the behaviour it evaluates. Once the target is met, there is still the question of whether the actual service improved and whether costs rose; if only the target figures flow back into the evaluation, both of these outcomes may be missed. Drawn for this book.
Four hours can change a hospital
England once faced the problem of emergency patients being held for long periods. A patient who had arrived at hospital might still wait and wait for assessment, a bed or the next arrangement. The government turned time into an explicit requirement: most emergency patients should be admitted, transferred or discharged within four hours. In 2005 the target demanded that 98 per cent be dealt with in that time.
It is not enough for the emergency department alone to move faster. If a patient who needs admission has no bed, however hard the frontline clinicians work, the next step cannot be completed. The hospital has to manage the coordination of tests, beds and the various departments; waits that had easily been seen as internal to emergency medicine began to become the responsibility of the whole hospital.
Researchers interviewed emergency department leaders at nine hospitals in 2008 and published their results in 2011. They recorded the changes in coordination, resources and process that interviewees reported, and also the pressure the time target brought and the doubts about quality.69 These interviews help us understand how hospitals responded to the target; whether patients received better care as a result still has to be checked against the corresponding clinical outcomes.
If decisions cluster just before the four-hour deadline, the patients’ circumstances need further examination. Some may be tests, beds or transfers that had previously been delayed and were at last arranged in time; others may have been hurried out while they still needed observation. The distribution of times can point to where investigation is worthwhile; only the clinical picture, the subsequent outcomes and the actual process can help tell the causes apart.
Setting a target sometimes brings people who had each been busy on their own to work together on a single difficulty. But how the difficulty is named also decides who can join that work.
Consider a hypothetical job-help form. It asks: “What is the main difficulty you face in looking for work?” It offers only three options: no work experience, not knowing how to write a CV, and not knowing how to find vacancies. There is no box for anything further.
Someone may lack none of the three; the real constraint is a clash between the time needed to care for a family member and the fixed shifts the vacancies demand. What needs comparing here is two ways of keeping the record:
| Replies the form allows | Whether the record can retain this difficulty |
|---|---|
| Only one of the original three options | The clash between shift patterns and caring responsibilities cannot be recorded faithfully as a different kind of cause |
| Answers not on the list are allowed, and a procedure reviews the categories | There is a chance of recognising situations the original options left out; it still has to be checked whether the additional information actually enters the handling of the case |
In the first format, the omission in the classification cannot remove itself by collecting more copies of the same form. The person filling it in may give up, pick an option reluctantly, or look for another channel to explain; which of these happens needs actual investigation. What can be settled first is that the original options provided no place where this difference could be faithfully written down.
Then ask a further question: what if the help on offer is also allocated according to those three options alone? At that point the form is no longer merely describing difficulties. It may affect which kinds of help can be obtained, and it gives people reason to adjust what they say to fit what the institution will accept.
If the answers received match the three options more and more closely, that cannot be taken directly as evidence that the classification is correct. People may have learned how to answer in order to stand a chance of getting help.
Revising the form is a start. Suppose the additional box at last lets the person explain their caring responsibilities, and the caseworker, having read it, confirms it with them; the caseworker then knows that another CV course will not solve this difficulty. What has to be looked for next may be a working arrangement that fits, or other support for caring. Whether such help can be provided already runs up against the limits of how resources were originally allocated.
A new answer may call for a new practice, and it may call for deciding afresh whom this service is meant to help and how far it should go. Folding the exception back into the nearest old option keeps the reports tidy; hearing the exception out may let managers discover that the original question was asked too narrowly.
After a manual is revised
The Diagnostic and Statistical Manual of Mental Disorders, published by the American Psychiatric Association, is usually shortened to DSM. It organises diagnostic names and criteria so that research and clinical work can communicate in a reasonably consistent language. Medical, insurance and educational arrangements in different places also refer to diagnoses, so the classifications in the manual may, through these various institutions, affect the resources a person can obtain.
In 1980 the third edition formally included post-traumatic stress disorder. People living with the long-term effects of trauma had existed before that, and had been described and treated under other names. The new diagnosis provided a shared name and shared criteria, allowing researchers to compare cases in a more consistent way, and giving clinicians and the people concerned one more way to explain a difficulty and seek help.70 Whether resources actually increased still depended on how subsequent services and institutions adopted the classification.
In 1973 the Association decided to stop listing homosexuality itself as a mental disorder. That change involved research evidence, professional dispute, action by the people concerned, and a reassessment of which conditions should count as illness. The related diagnostic names continued to be adjusted afterwards, and stigma did not vanish with a single revision; but changing the diagnostic status put the reasons that had used a medical classification to support certain kinds of treatment under direct challenge.
The fifth edition, in 2013, brought several previously separate diagnoses, Asperger’s syndrome among them, into autism spectrum disorder. The Association explained that some of the old categories had not been applied consistently across different clinical settings, and that there were research and diagnostic reasons for moving to a more integrated description. For people who had come to understand themselves under an old name, this also touched community, identity and the way they explained their own experience to others. How the criteria are revised, and whether a person still wishes to use a particular name for who they are, are not entirely the same decision.
The philosopher Ian Hacking called the back-and-forth influence between classifications and people “looping effects”: a classification changes how people are treated and may change how they understand themselves; and as people accept, reject or adapt these names, researchers in turn rethink the classification.71 To study such effects, one has to trace which services, expectations and actions a diagnosis actually changed, and not infer every person’s life from the name alone.
A child who receives a diagnosis may gain support that had been missing, and may also have their abilities underestimated. These two consequences need to be checked separately. Reducing stigma can proceed alongside taking the child’s difficulties seriously; keeping a useful diagnosis still leaves room to change expectations that are set too low or services that do not fit.
The people an average score cannot see
Suppose a tool gets ninety-five of a hundred cases right. Knowing that figure alone, we do not yet know where the five errors fall: scattered across all kinds of situations, or almost all in the same predicament? Nor do we know whether the consequence is a notice that can be corrected, or the loss of an opportunity that is hard to recover.
Judging someone who does not meet the conditions to be eligible, and keeping out someone who does, sometimes need different channels before they can be discovered. The former enters the process, and their later performance may leave a record; the latter never gets the opportunity at all, and the system may never see them again. Follow-up data on those already admitted cannot by itself prove that everyone excluded ought to have been excluded.
Appeals and additional information can bring back errors that were not seen the first time. If someone wrongly excluded has a way to submit further evidence, for instance, the institution has a chance to look at their eligibility again. It also has to be checked whether the appeal takes so long that it cannot be used, whether the information needed was obtained, and whether the person receiving the appeal has the authority to correct the decision. A procedure listed on paper is not enough to show that errors can be dealt with.
A frontline worker may notice a kind of exception very early, yet have no field in which to record it; the person who has a field may only be able to add a note, without being able to change the decision; and the person with the authority to change the rules may see only the aggregated figures.
The form’s additional explanation may run into an obstacle here too. The caseworker sees the clash between shifts and caring; those above receive only the statistics for the three established difficulties. Only if the original explanation is sent along to the people who can adjust the service does the new problem stand a chance of changing how resources are allocated.
Once it arrives, someone still has to respond. A decision that a service cannot be added for now also has to state which resources or which authority constrain it; only then can the records accumulate into a fresh discussion when the same kind of need appears again. The person concerned should not have to prove from scratch, every time they are passed to a different caseworker, that the same difficulty in their life really exists.
Figure 14.2 How the people who know about a problem get it dealt with. The right-hand diagram sends the original records and additional explanations to the people with the authority to act, and returns the decision to the frontline; the dotted lines are the added routes. This is a hypothetical comparison of institutional design, not an organisation chart of Wells Fargo. Drawn for this book.
A worker who sees a failure still may not be able to decide alone what a new rule will mean for everyone. The person who holds the authority to decide may not understand the ground any better than the worker does. Bringing raw experience into the discussion does not require handing all authority to one side first; it requires that each side’s reasons can actually change the proposal, and that the person responsible for the decision then explains which arrangement has been adopted.
Some results only show later. After the service is adjusted, whether the person concerned really has a better chance of finding work, and who has been given a heavier load by the new arrangement, both need new reports. That is how an institution can learn what its original classification never taught it.
But this depends on the people who raise the problem not being pushed out for making things more troublesome. An institution that lets only its designers name failure will find it hard to learn from experiences its designers never had. This limits knowledge, and it limits participants’ influence on their shared life.
Placing ability outside the individual
Work distributed across different positions can also make up an ability that no individual had to begin with. As an aircraft prepares to land, the speeds that need attention at different stages depend on conditions such as weight and flap setting. The pilot has to know more than a single number; in the midst of a busy operation, they also have to recognise how the current speed relates to the speed needed next.
Studying the cockpit, Edwin Hutchins observed that reference cards, spoken confirmations, instruments and speed bugs all take part in this work. The required speeds can be looked up in advance and placed on the instrument as markers; afterwards, looking at where the needle stands in relation to the markers completes part of what might otherwise have required remembering the numbers and then comparing values.72
The markers move some of the work to beforehand, and they also turn a comparison of values into a comparison of positions in front of the eyes. If they are set wrongly, they preserve the error just as faithfully; the confirmation between the pilot and their colleagues is part of this ability. Testing one person alone on how much they can remember without tools would miss what the actual operation relies on.
Andy Clark and David Chalmers’s “The Extended Mind” makes a stronger philosophical claim: under suitable conditions of coupling, external resources can become part of a cognitive process, and even of certain beliefs. The claim concerns how the boundary of the mind is to be drawn, and it remains a matter of philosophical dispute.73
The authors use a thought experiment to make the question concrete. Two people want to go to a museum; one recalls the address from memory, while the other, because of difficulties with memory, has long kept important information in a notebook carried everywhere, and turns to it as a matter of course when needed. If the notebook is reliably available, its contents consistently trusted, and it takes a continuing part in the person’s life, should we exclude it from the relevant cognition simply because it lies outside the skin?
Even without ruling for now on whether the notebook belongs to the mind, we can still study how it takes part in a person’s actual abilities. The same holds for the cockpit: the reference cards, the markers and the colleagues’ confirmations are all factors that have to be examined when explaining performance and error.
As the scope of cooperation widens further, new problems appear. Departments in a company can exchange data yet pursue different interests; schools can share student records yet hold different ideas of what good education is. Doing cognitive work together does not mean that everyone has the same purpose or a single shared consciousness. In analysing an organisation, the disagreements between participants, and what each of them is able to decide, still have to be kept in view.
Cooperation between disciplines also meets institutional differences of this kind. Different disciplines may genuinely need different concepts, instruments and training; within a university, they also organise their own courses, award their own degrees, and use their own systems of publication and review. Researchers who want to work together therefore face more than conceptual problems.
For instance, a study that needs the methods of two fields at once may first have to settle: which side can assess the other side’s evidence? To which kind of reviewer should the results be submitted? Is the time participants put in recognised within each side’s system of evaluation? These institutional arrangements affect whether the cooperation can continue. If the whole difficulty is put down to “the two sides think differently”, there is no training or evaluation system left that could be changed.
Researchers can first spend time learning to understand the other side’s evidence, and they may also find that what is hardest to change is the evaluation system. Two people already know how to work together; the units they belong to may not give them the conditions to keep going.
What questions a talking system brings back
AI can learn from the text, images and other data that people have accumulated over a long time, and can also acquire abilities through different kinds of training and tools. Some work that once required a person to learn for many years is work in which machines can now provide useful results; how ability is distributed, though, cannot be guessed by following the order in which humans grow up. Doing a hard task well does not automatically make every apparently simpler task reliable.
Before discussing “whether it is intelligent”, one can first sort out which question one wants to ask. The same answer may concern function, degree of resemblance to humans, internal mechanism or subjective experience:
| What you actually want to ask | Where to start checking | What this does not answer on its own |
|---|---|---|
| Whether it can complete a given piece of work | The task, its conditions, stability and failures | Whether it completes the work in the same way a person does |
| Whether its behaviour resembles a person’s | Language, actions and modes of interaction | Whether it has the same feelings |
| What usable representations have formed inside it | How information affects its operation and transfer | Whether this is enough to produce subjective experience |
| Whether it has subjective experience | First clarify what experience means and what evidence would be acceptable | A single human-like answer cannot close the case |
Being able to produce sentences about grief is an observable performance; whether it feels grief still needs other arguments. That it uses statistical methods internally is likewise not enough on its own to decide whether it can understand. Only once the questions are sorted out do we know which kind of claim a given observation supports.
Deployment also makes a model part of other people’s environment. A classification that exists only on paper and one that takes part in a great many screenings every day have different consequences. Users adjust their inputs, the people being classified respond to the rules, and later data may be affected by earlier decisions. So the original test scores still matter, yet they are not enough to answer how the new behaviour, power and costs are distributed.
In a hypothetical workflow, fixed programs can check clearly defined formats and constraints, AI or other search methods propose candidate solutions, and suitable trials then compare the results. AI might also first judge whether the conditions for a known method are met and then call a fixed procedure; if the results turn out anomalous, another approach is tried. Whether this division of labour is useful needs to be confirmed by testing on the task.
Some low-impact operations can be carried out first and logged, to be checked afterwards; some risks need to be constrained in advance. When a person takes over, they also have to be given the original situation and enough time. Which party seems more human, and which is called the intelligent core, cannot substitute for these arrangements. As for whether to change the shared goal, that still requires the participation of those who bear the responsibility and those who are affected.
Each stage may also receive the errors of the stage before. A model’s explanations need to be checked against the original data, the tool results and the operation logs; if the generator and the checker would miss the same kind of problem, running it once more does not necessarily add much capacity to catch errors. When designing a check, state which kinds of error it can find.
A person with only a few seconds left, and no view of the crucial information, can hardly carry all the checking that the process on paper hands to them. Who is actually able to spot a problem, and who has the authority to halt or change things, should be designed together with the allocation of responsibility.
Self-correction may serve only the original goal
A system can be very good at correcting errors while allowing only one kind of error to be named. If a service system asks only whether processing time has been shortened, it can keep learning to close cases faster without ever learning to recognise whether the person’s problem has actually been dealt with. Even if feedback from appeals is added, so long as appeals are still defined as a burden to be cleared, the feedback may help the system exclude dissatisfaction more efficiently.
Where the system once counted only the time taken to close a case, it now has to trace where that person’s difficulty went after the case was closed. Receiving a new piece of experience sometimes makes us rethink even what “done” means.
Those who propose a new goal also need to state their reasons and the likely costs; others may still disagree. But if the institution accepts only suggestions about “how to close cases faster”, and never lets anyone discuss whether a case has been resolved, the part that most needs changing stays permanently outside the discussion.
15 — What Makes an Understanding Worth Relying On
When we make a judgement, we are also changing how the next judgement can be made. Relying on a tool may save effort today while letting a skill that goes long unused slowly grow rusty; keeping a record of one failure adds work now, yet those who come later may be spared a stretch of wasted road because of it. The ways we understand the world are themselves changing the capacities and conditions within it.
Begin by comparing two hypothetical public service systems. At present they handle the same cases, and on the results that have been checked they are equally accurate. System A keeps suitable records of its failures, lets those affected add circumstances the original classification had no room for, and has someone responsible for dealing with errors once they are confirmed. System B gradually deletes its failure data, restricts outside comparison, and makes it harder for the people who raise exceptions to go on taking part.
Judged by the current rate of correct answers alone, there is no difference between them. But B is reducing the means by which we could learn about its later performance. Without records, it is hard to compare whether an error keeps recurring; without access to other sources, it is hard to know whether the original results were complete; once the people who raised exceptions have left, the omissions in the classification also become harder to see.
B may still get the next case right. What is disappearing is the basis on which anyone could later judge whether it did, and the means of making a correction once a problem has been confirmed. A choice made today has already altered the conditions for knowing in the future.
Personal candour cannot maintain these capacities on its own. Someone willing to admit mistakes may still not know where they went wrong if the only cases they can read about are the successes; an institution that speaks with confidence, if it submits to independent checks that have real effect, may be easier to examine than one that is modest in speech yet refuses to hand over its data. Attitude has to be matched by data, skill and authority before it can actually change a judgement.
Why should effort be spent on preserving the capacity to catch errors? A value judgement needs to be stated openly here: when an error could cause serious loss, and the chance of avoiding it can be kept open at a proportionate cost, we have reason to make that upkeep part of our choice. This is no command derived from the bare fact that people make mistakes. It gives weight to avoidable loss, and at the same time it requires costs, rights and other needs to be weighed against one another.
From this, the book puts forward a claim for evaluating long-term reliance: while serious uncertainty remains, the choice of a method should also take into account how it affects the capacity, later on, to obtain evidence, to learn, and to correct serious errors.
What is being evaluated here is the reason for continued use. The correct answers a book already contains do not become wrong because its author refuses criticism; but on content not yet verified, on new situations of use, or on future editions, a refusal to have errors checked affects what we have to go on in continuing to believe it. Whether an answer is true, what evidence there already is, and how evidence will be obtained and corrected later are three things that need to be set out separately.
The distinction also changes how each case is handled. A method that performs well yet lacks any channel for checking may be helped simply by adding external verification first, without the method itself having to be rewritten straight away. A method whose procedures are open yet which keeps giving wrong answers may need to be revised, or even relieved for a time of some important task; a willingness to be checked is not the same as already being able to do the job. Between A and B, the actual costs still have to be compared, and it cannot be concluded that A should be adopted whatever the price simply because it keeps more procedures in place.
Figure 15.1 Present performance is the same; the later capacity to catch errors may already differ. Movement to the left marks weakening conditions for catching and correcting errors, and is no prediction that the next case will certainly go wrong; the figure has no measured scale. Drawn for this book.
Conversely, adding an appeal channel that can actually overturn a determination, even before it has raised the average score, may allow people who previously had nowhere to appeal to point out omissions in the classification. Its effect still has to be checked and its cost assessed; but it cannot be concluded that the change has no value simply because the short-term score has not risen. What long-term reliance requires us to compare also includes which evidence can be obtained in future, and who is in a position to make that evidence change a decision.
Catching errors has costs too
Between services A and B, keeping more records is not always better either. More checking may lower efficiency, and it may intrude on privacy. Publishing certain data in full may expose vulnerable people to fresh harm. Allowing every objection to halt the work may leave an important service unable to run at all.
Preserving the capacity to catch errors has to be weighed together with cost, time limits and rights. Some information can be seen by independent examiners bound by obligations, without being opened to everyone; some exceptions are better reviewed at regular intervals than allowed to interrupt the work each time. The concrete arrangements have to be designed separately.
If an error-checking arrangement costs a great deal, adds almost nothing to what can actually be detected, or causes a graver loss that cannot reasonably be accepted, there may be no sufficient reason to keep it. Being given the name of oversight does not exempt it from a check on its costs.
There are also situations in which going on checking the same thing is simply not worth it. If we truly had a sufficient guarantee that a system must be correct under every relevant condition, and that it would never stray beyond those conditions, then checking the same thing over and over might do nothing except add cost. This is why the claim above was limited to cases where serious uncertainty remains and reliance will continue; for a result that is already guaranteed, the mere fact that it could be checked again is no ground for demanding more resources.
To apply such a guarantee to the system in front of us, its scope has to be established. Which premises does the formal proof use? Which inputs do the tests cover? Will actual use go beyond those conditions? So long as some part that could affect the outcome remains unconfirmed, the reason for catching errors has not entirely gone.
An opaque interior does not mean that every kind of reliable evidence is lacking. Performance records, external measurement and suitable comparison can sometimes supply sufficient reason for use. Whether the internal steps need to be understood depends on whether doing so would settle an important question still unconfirmed; transparency in itself cannot be treated as a necessary condition for all knowledge.
Letting those who come after do more
Keeping data is useful because someone can take it and compare afresh; writing out reasons is useful because someone can learn them and point to a step that does not hold. If the data exist but no one has permission to read them, if the method is written down but practice and questioning are never allowed, what gets handed down may be nothing more than the name of an authority.
What we need to build together therefore includes data that can be obtained, practices that can be learnt, and openings through which new findings can enter decisions. These support one another. A new measurement may expose the inadequacy of an old classification; the account of someone directly affected may point to what ought to be measured; reasons made public allow people far away to redo the comparison. No single person has to know everything first, and other people’s results can still become one’s own new capacity.
More people able to take part does not make error disappear. A newly added examiner may also be swayed by interests, or may simply carry over the judgements of the original group. What is worth comparing is how much more the data and practices they bring can actually uncover.
The value of an added check can be relative and limited. Two checks that fail in different ways may expose problems better than one check repeated; making the source of the data traceable may be easier to verify than an authoritative signature alone; separating the authority to decide from certain interests may reduce particular distortions.
These comparisons do not have to wait for a final examiner who never errs. What they need are concrete reasons showing how the added arrangement changes risks already identified, and whether it is worth its cost. If a layer of oversight merely reproduces the blind spots already present, this book has no reason to count the added layer as progress.
Snow in Chapter 2 completed his comparison by relying on residents, landlords and water-supply records; the reader in Chapter 10 followed the reasons and learnt a method they had not known before; the service in Chapter 14 reconsidered what it offered, starting from difficulties the original form could not record. Together these pieces of work enlarged what later people could ask, check and do. Their value goes beyond making the old answers wrong a few times less often.
A coherent way of thinking can form through learning of this kind. Faced with different problems, it lets proof, interviewing, sensory practice or negotiation each do their own work; choosing among them in turn requires reasons that can be stated and that bear on the problem. One’s own experience submits to the same demand. Even “this is how I usually check” can turn out to be a practice in need of revision.
This direction carries forward the thinking introduced earlier. Peirce and Dewey placed inquiry within difficulty, action and experience; bounded rationality brought time and ability into the explanation of judgement; Hutchins’s research, together with Clark and Chalmers’s argument for the extended mind, made arrangements beyond the individual harder to ignore. Putting these threads together, we can press the question: will the capacities we borrow today give the people of tomorrow a better chance to understand and change their own situation, or make it ever harder to see where something has gone wrong?
After seeing clearly, a commitment still has to be made
Yet those who come after becoming ever more able to achieve some purpose may also be the more worrying prospect. An oppressive institution can learn to recognise its own failures and maintain the oppression more effectively. Understanding how to get a thing done has not yet answered whether it is worth doing. Effectiveness in knowing cannot by itself manufacture ethical legitimacy. Whether those affected can make demands, and whether they have a right to be free of certain treatment, involves a normative position that has to be defended directly.
The position this book takes is that people should have a real opportunity to offer reasons, raise questions and seek suitable redress concerning arrangements that deeply affect their lives. How neural reflexes and control systems work has not chosen this position for us; a commitment to the standing of persons and to living together has to be defended with ethical and political reasons.
If the service system of Chapter 14 wants only to close cases faster, appeals may come to be treated as a number to be kept down. Giving those affected the right to question whether “closed” means the problem has been solved changes what the system ought to learn. Even so, once every party is able to give reasons, a choice everyone agrees on will not necessarily emerge.
In 2009, the breast cancer screening recommendation published by the U.S. Preventive Services Task Force provoked controversy. The recommendation of the time left the question of whether to begin regular mammography between the ages of forty and forty-nine as a decision that had to take account of individual circumstances and the weight given to benefits and harms; for ages fifty to seventy-four, it recommended screening every two years. What is discussed here is a historical document from that year, and it cannot serve as personal screening advice today.74
Screening looks, among people without relevant symptoms, for signs that may call for further examination. Finding certain cancers early can bring benefit, yet an abnormal result may in the end prove not to be cancer; and some cancers that are found and treated might never have caused symptoms or death within that person’s lifetime. These different situations mean that “a few more cases found” cannot by itself answer the question of overall benefit and harm.
When comparing effects, the denominator also has to be known. Take first a set of hypothetical figures unrelated to any breast cancer screening data: some programme reduces a risk from four people in every hundred to two. To say “the risk is halved” is correct, and to say “two fewer people in every hundred” is also correct; the second lets the reader know at the same time how large the original risk was. When only the proportional reduction is reported, people readily form different pictures of the actual difference.
The controversy of that year concerned whether the evidence was sufficient, how the models estimated the effects, and how benefits and harms should be expressed. Even where some effects are estimated more clearly, it remains to be weighed how much expected benefit is worth how much additional examination, anxiety and cost of treatment. Where an individual has to decide, they also have to be given explanation and support they can understand; simply handing them the choice does nothing to lessen the difficulty.
Richard Rudner proposed in 1953 that how strong the evidence must be before a hypothesis is accepted may depend on the consequences of judging wrongly. Heather Douglas later went further, studying how such inductive risk bears on research methods, the reading of data and inference.75 A claim of this sort brings “how great a risk of misjudgement to bear” into the discussion; the evidence itself must still be handled truthfully.
For example, the grave consequences of failing to detect a harmful substance can lead us to demand different conditions of testing; they cannot lead us to rewrite a measured value because we hope the substance is harmless. That values take part in how uncertainty is borne does not mean the facts can be shaped at will.
Admitting that one may have seen wrongly does not require standing forever in the middle of every dispute. You can be willing to revise your judgement of an institution’s effects while making clear that you will not support shifting heavy costs onto people who have no choice. The first concerns evidence; the second puts forward a position that has to be defended.
That position has a history of its own. Family, education and experience take part in what we care about, but finding the causes that formed a position is not the same as proving it wrong. What still has to be asked is this: knowing all that, are we now willing to go on owning it, can we explain our reasons to those it affects, and are we willing to face the consequences it brings about?
Reason, feeling and experience can all lead us to reassess a value. Coming to understand another person’s situation may change the costs we were previously willing to accept; establishing the facts may take the ground from under an earlier anger. If we go further and want to turn our own commitment into a shared rule binding on others, we must also explain who can take part in the decision, and why those affected should accept such an arrangement.
From understanding back to action, and back again
The choosing, checking and practising discussed so far move in two kinds of loop. To get something clear, we form an interpretation from observation, state an expectation, and then revise it with new material. To get something done, we choose an action with a purpose in mind, and when the results arrive we check the method and may also change the original purpose.
The two loops can interleave. A trial run under suitable conditions can help identify a need that could not be articulated before; a research finding may strip an otherwise attractive goal of its justification. But action is not an all-purpose licence to experiment. When someone will be hurt, when rights and interests will change, when an opportunity may not be recoverable, a thing cannot be carried through simply because it would aid learning.
A single check may also change what was originally to be done. Recovering the record made on the spot may mean correcting only one number; hearing out the experience of someone never asked before may reveal that the scope of the problem needs to widen. There is no need at that point to run through every earlier chapter again; following the difficulty that has actually appeared is enough.
The stopping question of Chapter 11 still holds: what might the next step of analysis change, and how much time is it worth? Even so, that a question will not change the action in front of us for now does not mean it is never worth studying. You can complete the present decision first and set aside other time for the needs of learning and understanding for their own sake. The two pieces of work do not have to be settled on the same evening.
With time, some distinctions become second nature. You do not have to recite the definitions of observation and interpretation before noticing that an accusation carries an unproven guess inside it; nor do you have to draw the whole institutional map before knowing that the person in front of you lacks the authority to deal with a problem they have already seen. Once a method has been learnt, it can take up less of one’s awareness, leaving attention for the differences that belong to this occasion alone.
One way of thinking, with room for different practices
The sentence from the preface, “I have thought about it carefully”, now has a continuation. What information was to hand while thinking, whose methods were borrowed, what still remains undealt with? Only in answering these questions do we learn what that care accomplished.
Reflection itself has to submit to this kind of checking. Going back over one’s reasons may turn up a contradiction, or it may amount to turning the same thoughts over and over, eating up time for observation, rest and action. If, because we advocate reflection, we no longer allow its effects and costs to be compared, we have made an exception for precisely the method we trust most.
Coherence has practical content here. If I hold that the origins of a view may shape other people’s judgements, I cannot declare that a statement has no formation history worth studying merely because it came from me. If I say direct experience deserves weight, I must also allow other people’s experience to raise difficulties for my classifications. The same reason can carry different weight on different occasions, but the difference has to be supported by the circumstances and cannot be settled by “this time it is me”.
A familiar operation can rely on practice that has already received feedback; an unfamiliar choice may call for looking up more data; in an emergency, one must use whatever capacity is most reliable at that moment. These differences need not break the coherence; what has to be accounted for is why time, evidence and consequences led us to switch to another practice.
Consistency is still not enough to guarantee correctness. A whole set of ideas may rest on wrong data, or may never have counted in the needs of some group of people. So besides checking whether one’s own statements are compatible with one another, they have to be compared against actual results, other people’s experience and objections.
Figure 15.2 Sometimes the answer has to change; sometimes the way of judging has to be reconsidered. The outer loop shows that the existing arrangement can be re-examined, without requiring the whole way of thinking to be rebuilt for every action. Drawn for this book.
To demand, every time, that the criteria be proved infallible before one is permitted to begin is to fall back into the difficulty of Chapter 11. Judgement moves forward on the knowledge it already has, and leaves places where one can come back later and correct it.
Those who come after will do more than correct us. Given the reasons and the chances to practise that we leave behind, they may ask questions we have not yet thought of. Making that possible is one more reason an understanding deserves to be handed down.
This book is in it too
This book has chosen certain stories, materials and ways of putting things, and so it readily draws the reader’s attention to some matters while passing over others. It spends many pages on method, but what a person sometimes truly lacks is time to rest, or someone willing to hear them tell the whole story through. To write all of that up as a failure of thinking would misread their situation just as badly.
If a reader finds that a distinction does not help, that an example does not match the data, or that following the analysis caused them to miss what ought to have been done, this book too needs to change. It cannot earn the standing of something worth relying on in advance, merely by advocating correction.
Afterword — After Understanding, Living Together
When the priest handed the child to the woodcutter, the killing in the wood still had no answer. Nor did we find one for him in the chapters that followed. Whether the woodcutter stole the dagger, and how he will care for the child from now on, remain things that still have to be understood.
Yet the two men had already responded from within that incomplete understanding. The story did not wait for everything to be established before letting someone carry the child away in his arms. This has kept me attentive to one thing. We want to get matters clear, and often we want it because other people live inside those matters with us, waiting for an answer, or bearing the consequences of there being none yet.
Knowing this does not suddenly make judgement easy. Some harms have to be stopped at once, while some causes take a long time to trace; some people can state their reasons clearly, while others have not yet found the right words. We need knowledge, and we also need the time, the tools and the help from one another that let knowledge arrive.
To write down a reason is to hand one’s own understanding to someone else. They may carry on using it, or they may point to what we failed to see. Someone is willing to leave a record of a failure; someone is willing to demonstrate a movement that cannot be put into words; someone brings into the discussion an experience that previously had nowhere to be voiced. Later understanding begins from concrete work of this kind.
We cannot decide for the next person what they ought to see. What we can do is make sure that today’s understanding, useful as it already is, does not become the reason they can no longer ask tomorrow.
Coordinates and Sources
The historical facts, studies, interviews and retellings of literary works in the main text rest on the sources listed below. Images are credited separately in their captions and image-source notes; historical image files are kept apart from the diagrams drawn for this book. Literary works are kept for their narrative and intellectual value; interviews are kept as the experience of those interviewed and as reconstruction after the event. Cases explicitly marked as hypothetical are used to test reasoning and do not pose as actual events. The evaluative claims and institutional proposals in Chapter 15 are positions this book puts forward and defends; they should not be read as conclusions that the works below jointly endorse.
Emily Pronin and Matthew B. Kugler (2007). “Valuing thoughts, ignoring behavior: The introspection illusion as a source of the bias blind spot.” Journal of Experimental Social Psychology 43: 565–578. Original paper. The study compares the different weight given to introspection, behaviour and other material when assessing one’s own biases and those of others; it does not mean that every person in every judgement is necessarily like this, and whether a reason holds cannot be decided directly from how it came about. ↩︎
Heinz Wimmer and Josef Perner (1983). “Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.” Cognition 13(1): 103–128. Original paper. The ball, the box and the drawer in the Preface are an illustration adapted from the false-belief task based on a change in an object’s location, not a verbatim retelling of the original experimental scene. The book uses it to show how what different people know can be understood within one scene; passing one task is not equated with full self-awareness, and it is not claimed that everyone has the same moment of insight. What this book calls “the second realisation” is a metaphor the author uses to describe counting oneself in, not a further universal developmental stage confirmed by this research. Nor does the children’s task on its own establish the normative demands this book makes of adult judgement. ↩︎
Oskar Pfungst (1911). Clever Hans (The Horse of Mr. von Osten): A Contribution to Experimental Animal and Human Psychology. Translated by Carl L. Rahn. Henry Holt and Company; the German original was published in 1907. Full text of the original study. The whispered addition in Chapter 1 follows the book’s experiments in which “the questioner does not know the answer”: 3 correct out of 31 tests with unknown sums, 29 correct out of 31 tests with known sums. For the postural and head signals see the book’s analysis of movements; these figures are limited to that set of tests. ↩︎
Immanuel Kant, Critique of Pure Reason (1781/1787). Gutenberg English translation. Chapters 1 and 7 use it as the intellectual background on sensibility, concepts and the status of geometry; the bearing of non-Euclidean geometry on Kant’s thought is still disputed among interpreters, and a brief case is not used to declare his whole epistemology void. The chapter discusses the problem that non-Euclidean geometry poses for Kant’s argument about geometry, and does not use a spherical diagram or a single mathematical development to declare his whole epistemology overturned. ↩︎ ↩︎2
Elizabeth F. Loftus & John C. Palmer (1974). “Reconstruction of Automobile Destruction: An Example of the Interaction Between Language and Memory.” Journal of Verbal Learning and Verbal Behavior, 13, 585–589. Original paper. Chapter 1 keeps the two experiments separate, and does not write the speed estimates and the broken-glass question a week later as one test on the same group of people. The chapter’s claim is limited to the difference between the question asked and the later answer that the study showed; it does not infer that memory can be rewritten at will, or that everyone is affected in the same way in every situation. ↩︎
John Snow, the map recording the 1854 outbreak in On the Mode of Communication of Cholera. Image file and rights information. The original is in the public domain; this book uses the 3,840-pixel version provided with the file, with the content of the map unaltered. The clustering of deaths on the map cannot on its own replace the investigation of the water source. ↩︎
John Snow (1855). On the Mode of Communication of Cholera, 2nd ed. John Churchill. Original text of the Broad Street investigation, original text of the water-company comparison, hosted by the UCLA John Snow site. Chapter 2 retells the workhouse, the brewery, the Hampstead case and the water-supply investigation from the book. The figure of 535 refers to the workhouse’s resident population; the eight- to ninefold figure compares deaths per 10,000 supplied houses in the first seven weeks of the outbreak, not the risk per 10,000 people. The book also records supplies not yet identified and cases that were estimated. ↩︎ ↩︎2
Philip E. Tetlock (2005). Expert Political Judgment: How Good Is It? How Can We Know? Princeton University Press. Digital copy of the book. Chapter 2 adopts the research design of long-term tracking and comparison of forecasts; the weather figures for calibration and discrimination are this book’s own example. The chapter does not use one general comparison to pass judgement on all experts, nor does it treat calibration as wholly unrelated to intelligence or professional knowledge; the research should be understood by its questions, its baselines and its scoring. ↩︎
Lewis Carroll (1865). Alice’s Adventures in Wonderland, Chapter III, “A Caucus-Race and a Long Tale”. Original text. Chapter 3 retells the episode in this book’s own words; the reading of purpose, evaluation and the allocation of costs is this book’s argument. ↩︎
Underground Railways of London (1908). Image file and rights information. The file page lists the author as unknown and the work as public domain; the 1,932 × 1,530 pixel image supplied is used. It serves as a historical map from before Beck’s design and is not treated as a controlled comparison of the same network. ↩︎
Henry Charles Beck, London Underground Transport (1933), the second edition of that year; collection number 8727.003. David Rumsey collection record and licence details. Image credit: David Rumsey Map Collection, David Rumsey Map Center, Stanford Libraries. The image is used under the CC BY-NC-SA 3.0 licence listed by the collection; it has not been cropped, recoloured or redrawn. This licence for the digital file applies separately to the right-hand image of Figure 4.1. ↩︎
Transport for London Corporate Archives. Research Guide No. 24: Harry Beck. Archive research guide; London Transport Museum, the 1933 pocket Underground map in the collection. Chapter 4 keeps the history of the design and its adoption, and does not merge the payments for different commissions into a single anecdote. ↩︎
Jill H. Larkin & Herbert A. Simon (1987). “Why a Diagram Is (Sometimes) Worth Ten Thousand Words.” Cognitive Science, 11(1), 65–100. DOI of the original paper. Chapter 4 carries forward the distinction between informational and computational equivalence; the multiplication of numbers and the route diagram are this book’s teaching constructions. ↩︎
Bureau d’Enquêtes et d’Analyses (2012). Final Report on the Accident on 1st June 2009 to the Airbus A330-203 Registered F-GZCP Operated by Air France, Flight AF 447 Rio de Janeiro–Paris. Investigation report, preserved by the FAA. For the sequence of control inputs and warnings see pages 21–24 of the report; for the analysis and conclusions see sections 2 and 3. Chapter 5 uses it to distinguish the control inputs, the validity of the airspeed readings, the state of the aircraft and the crew’s understanding; the problem of hand-over at the interface is the angle of analysis this book draws from it. ↩︎
Joel Spolsky (2002-11-11). “The Law of Leaky Abstractions.” The author’s original text. Chapter 5 keeps the name “leaky abstraction”, separating the empirical generalisation from engineering from the cross-domain argument this book makes on its own account; the example of two sets of data with the same summary where a new question demands different answers offers a necessity argument with explicit premises, and does not claim that one engineering article proves every abstraction must be wrong for every use. ↩︎
IETF (2022). RFC 9293: Transmission Control Protocol (TCP), in particular the description of the service in section 2.2 and the handling of connection failure in section 3.8.3. Formal specification. A reliable, ordered byte-stream service does not amount to delivery within a fixed time or a connection that never drops; the main text does not write failures the specification allows as violations of the specification by the protocol. ↩︎
Lisanne Bainbridge (1983). “Ironies of Automation.” Automatica, 19(6), 775–779. PDF of the original paper. Chapter 5 adopts its analysis of the manual tasks, the skills and the difficulty of taking over that remain after automation. ↩︎
Gregory A. Daddis (2011). No Sure Victory: Measuring U.S. Army Effectiveness and Progress in the Vietnam War. Oxford University Press. Publisher’s record. Chapter 5 uses the many performance metrics and the difficulty of reading them to discuss strategic judgement; this analysis does not amount to a single cause of the war’s outcome. ↩︎
NASA Science, Building Blocks. Chapter 5 uses it only to distinguish dark matter, dark energy and their observational background, and does not claim that the ultimate physical nature of either is known. ↩︎
Hans Christian Andersen, “The Little Match Girl”, in Andersen’s Fairy Tales. Gutenberg English translation. Chapter 6 retells the story in this book’s own words, keeping the sequence of the stove, the roast goose, the Christmas tree, the star and the grandmother, and the ending in which she is taken to God while what the street sees is a death. ↩︎
Bertram R. Forer (1949). “The Fallacy of Personal Validation: A Classroom Demonstration of Gullibility.” The Journal of Abnormal and Social Psychology, 44(1), 118–123. DOI of the original paper. Chapter 6 keeps the classroom demonstration, distinguishing the rating of the diagnostic instrument from the rating of the descriptive content, and does not let satisfaction stand in for the ability to discriminate. ↩︎
Ted J. Kaptchuk et al. (2010). “Placebos without Deception: A Randomized Controlled Trial in Irritable Bowel Syndrome.” PLOS ONE, 5(12), e15591. Full text of the study. Chapter 6 is limited to the study’s short-term, self-reported symptom outcomes and its research conditions. ↩︎
Thomas S. Kuhn (1957). The Copernican Revolution. Harvard University Press; NASA Science, Orbits and Kepler’s Laws. Chapter 6 keeps the comparison of calculation and theory in the history of astronomy apart from the basic account of modern orbits; it does not repeat the contrast of “degrees for the old tables, arcseconds for the new”, which comes without the object, the period and the precision conditions attached. ↩︎
Jon Kabat-Zinn (2003). “Mindfulness-Based Interventions in Context: Past, Present, and Future.” Clinical Psychology: Science and Practice, 10(2), 144–156. DOI of the original paper. Chapter 6 uses it for the formation of mindfulness-based stress reduction and its traditional background; it does not say that every ethical context has been removed by the course, nor does it claim in general that every use is effective. ↩︎
Euclid, Elements, the Greek text of the geometry with mathematical commentary; Girolamo Saccheri, Euclides ab omni naevo vindicatus (1733); Bernhard Riemann, “Über die Hypothesen, welche der Geometrie zu Grunde liegen” (lecture 1854, published 1868). The spherical diagram in Chapter 7 is a mathematical illustration, not a complete proof of the independence of the postulate. ↩︎
Neil Ashby (2003). “Relativity in the Global Positioning System.” Living Reviews in Relativity, 6, 1. Full text of the study. Chapter 7 adopts the explanation that comparing times requires relativistic correction; it does not treat one idealised conversion as a fixed daily positioning error for every receiver. The actual positioning error depends on the system’s configuration and corrections, and cannot be obtained by simply multiplying a time difference by the speed of light and treating the result as a fixed daily position offset borne by every user. ↩︎
Bertrand Russell (1912). The Problems of Philosophy, Chapter VI, “On Induction”. Original text. Chapter 7 keeps the chicken of the original and does not mix in the timing or the holiday of the later turkey version; this book’s practical discussion of extrapolation does not claim to have solved the whole problem of induction. ↩︎
IPCC (2021). Climate Change 2021: The Physical Science Basis, Chapter 1, Chapter 7. Chapter 7 distinguishes model ensembles, sensitivity ranges and emission scenarios. The assessed ranges carry their own confidence levels; the use of RCP8.5 does not automatically make it the most likely future, or a strict upper bound on all futures. ↩︎
Abraham Wald (1943; reprinted 1980). A Method of Estimating Plane Vulnerability Based on Damage of Survivors. Center for Naval Analyses; originally a set of memoranda for the Statistical Research Group at Columbia. Reprint in the archive, transcription of the text. Chapter 8 follows its research question, estimating loss probabilities from survivor data, and does not add dramatised dialogue between officers and mathematicians. The popular version, “add armour only where the bullet holes are fewest”, does not state the assumptions Wald’s estimate relied on. The chapter keeps the distinction between the probability of being hit and the probability of surviving a hit, and does not treat the count of holes as a complete rule for armour design in itself. ↩︎
Cardiac Arrhythmia Suppression Trial Investigators (1989). “Preliminary Report: Effect of Encainide and Flecainide on Mortality in a Randomized Trial of Arrhythmia Suppression after Myocardial Infarction.” New England Journal of Medicine, 321, 406–412. Record of the original paper, NHLBI study database description. Chapter 8 draws on the results of the specific treatment arms to discuss the difference between surrogate indicators and survival outcomes. ↩︎
F. W. Dyson, A. S. Eddington & C. Davidson (1920). “A Determination of the Deflection of Light by the Sun’s Gravitational Field, from Observations Made at the Total Eclipse of May 29, 1919.” Philosophical Transactions of the Royal Society A, 220, 291–333. Original paper. Chapter 8 keeps the expeditions, the measurements and the test of the theory apart, and does not add inner monologue for the observers. The chapter does not treat every plate as equally clear and fully consistent with the others, nor does it equate one measurement with the sole verdict on an entire theory. ↩︎
Karl R. Popper (1963). Conjectures and Refutations, Chapter 1; Popper (1978). “Natural Selection and the Emergence of Mind.” Dialectica, 32(3–4), 339–355. DOI of the 1978 article. Chapter 8 distinguishes the claim about testability, Popper’s own account of himself and his revision regarding natural selection. ↩︎ ↩︎2
Peter T. Boag & Peter R. Grant (1981). “Intense Natural Selection in a Population of Darwin’s Finches (Geospizinae) in the Galápagos.” Science, 214(4516), 82–85. DOI of the original paper. Chapter 8 uses the comparison of specific traits, food and survival to illustrate testable content, and does not treat one observation as the sole verdict on the whole theory of evolution. ↩︎
Albert A. Michelson & Edward W. Morley (1887). “On the Relative Motion of the Earth and the Luminiferous Æther.” Transcription of the original paper. Chapter 8 describes a fringe shift far smaller than expected; it does not write that the instrument produced no interference fringes, nor does it compress the whole later development of theory into a single overturning on the day. ↩︎
Ignaz Semmelweis (1861). Die Aetiologie, der Begriff und die Prophylaxis des Kindbettfiebers. Record of the book, with images and partial translation. Chapter 8 follows his record for the comparison of the clinics, the contamination hypothesis and the washing measures; the professional and emotional cost that admitting error may involve is raised in the main text as a general analysis, and is not used to pronounce on the motives of each historical figure. ↩︎
Akira Kurosawa (director), Shinobu Hashimoto and Akira Kurosawa (screenplay), Rashomon (1950), adapted from the work of Ryūnosuke Akutagawa. Criterion film record, Stephen Prince’s review and structural analysis, plot summary of the film. What Chapter 9 retells is the film: the shelter from the rain under the gate, the differing causes of death in the testimonies, the woodcutter’s changed account, the accusation over the dagger and the adoption of the child; the film’s framing plot does not belong to the content of Akutagawa’s original “In a Grove”. ↩︎
Thomas S. Kuhn (1977). The Essential Tension: Selected Studies in Scientific Tradition and Change. University of Chicago Press, the preface’s retrospective on reading Aristotle. Digital copy of the book. Chapter 9 treats it as the author’s retrospective, distinguishing understanding the meaning of the words, accepting a historical theory and the formation of the later work. The recollection is used to show how the use of concepts is reconstructed in reading; it does not mean that all of Aristotle’s claims in physics thereby hold, nor does it ascribe the formation of The Structure of Scientific Revolutions to a single moment of insight. ↩︎
Jonathan Haidt (2012). The Righteous Mind: Why Good People Are Divided by Politics and Religion. Pantheon, his own account of fieldwork in India in 1993. The author’s website for the book. Chapter 9 does not generalise a particular visitor’s experience into the position of a whole culture, and it keeps the different consequences that people of different standing may bear. ↩︎
Zhuangzi, “The Mountain Tree” and “The Secret of Caring for Life”. Chapter 9 retells the empty boat and Chapter 12 retells Cook Ding; the intellectual and literary context is kept, and this book’s contemporary extension is kept apart from the physiological research. ↩︎ ↩︎2
Charles Darwin (1887). Autobiography, edited by Francis Darwin in The Life and Letters of Charles Darwin. Text of the autobiography. Chapter 10 follows his retrospective account of how reading Malthus in 1838 combined with his earlier observations. For the background on population and the means of subsistence see Thomas Malthus, Chapter 1 of An Essay on the Principle of Population (1798). After forming his preliminary explanation, Darwin went on gathering evidence and developing the theory. ↩︎
Theorem Proving in Lean 4, Introduction and Axioms and Computation, official Lean documentation, checked on 2026-09-11. Formalisation requires that propositions be stated precisely, and proofs are checked against definitions, axioms and the rules of logic. Passing the check does not by itself guarantee that the formal model covers every relevant condition in reality. This book draws from it a discussion of the value of thinking that can be carried on by others; that extension is not a full philosophical claim made by the official documentation. ↩︎
Charles S. Peirce (1877). “The Fixation of Belief.” Popular Science Monthly, 12, 1–15. Original text. Chapters 10 and 15 use it to locate the intellectual source of the discussion of belief and inquiry, and do not use it to reduce pragmatism to “whatever is useful is true”. ↩︎
John Dewey (1910). How We Think. D. C. Heath. Full text of the 1910 edition. Chapters 10 and 15 carry forward its discussion of reflection, concrete difficulties and checking; the promise case and the bounded analysis steps in the main text are this book’s own construction. ↩︎
Herbert H. Clark and Susan E. Brennan (1991). “Grounding in Communication.” The authors’ public copy of the original chapter. Common ground and the mutual confirmation of understanding are updated as communication proceeds, and different media provide different conditions for this. The main text connects it to reading and the reconstruction of background, and does not equate the whole of empathic ability with this one model of communication. ↩︎
Tal Eyal, Mary Steffel, and Nicholas Epley (2018). “Perspective Mistaking: Accurately Understanding the Mind of Another Requires Getting Perspective, Not Taking Perspective.” Journal of Personality and Social Psychology 114(4): 547–571. Original paper. The paper reports twenty-five experiments; imagining another person’s perspective did not consistently improve the accuracy of judgement, while arrangements in which information was obtained through conversation did. Chapter 10 uses it to distinguish reconstructing thoughts from obtaining information; the conclusion is limited to the tasks tested and is not extended into a claim that empathy, imagination or historical understanding is generally ineffective. ↩︎
The Yijing (Zhouyi): the hexagram statements, the line statements and the Ten Wings; text of the Zhouyi. Chapter 10 keeps apart the combinations of symbols, the textual layers and the interpretive traditions. This book’s “using yin and yang as a way of asking questions” is a contemporary appropriation, not evidence of the predictive validity of divination. ↩︎
Gottfried Wilhelm Leibniz (1703). “Explication de l’arithmétique binaire, qui se sert des seuls caractères 0 et 1, avec des remarques sur son utilité, et sur ce qu’elle donne le sens des anciennes figures chinoises de Fohy.” English translation of the original. Chapter 10 distinguishes the binary arithmetic already developed, the hexagram diagram obtained later and the attribution of historical intent. ↩︎
Raymond Chen (2003). “What’s the Deal with Those Reserved Filenames Like NUL and CON?” An engineer’s historical explanation; Microsoft, Naming Files, Paths, and Namespaces. Chapter 10 distinguishes the historical explanation from the current rules of each interface; the documentation was checked on 8 September 2026. ↩︎
Koichi Yasuoka & Motoko Yasuoka (2011). “On the Prehistory of QWERTY.” Original study in the Kyoto University repository. Chapter 10 uses it to examine the popular story that the layout was “designed to slow typists down”, and does not declare one interpretation of the sources in it to be the undisputed and complete origin. ↩︎
U.S. General Accounting Office (2000). Year 2000 Computing Challenge: Lessons Learned Can Be Applied to Other Management Challenges. GAO/AIMD-00-290. Official report. Chapter 10 keeps the inventory, the remediation, the testing and the contingency planning, and does not adopt global total-cost figures that have not been separately verified. ↩︎
NASA, Apollo 11 Lunar Surface Journal, transcript of the descent and landing communications, in particular the report from mission time 102:38:26, the reply at 102:38:53 to continue the descent, and the repeated alarms that followed; Fred H. Martin (1994), a participant’s retrospective on the program alarms. The exchanges between the astronauts and Duke in Chapter 11 are retold from the transcript; the division of engineering work on the ground and the search for the cause after landing also draw on the editors’ notes and Martin’s later recollection. ↩︎
Herbert A. Simon (1978). “Rational Decision-Making in Business Organizations.” Nobel lecture and text; Simon (1990). “Invariants of Human Behavior.” Annual Review of Psychology, 41, 1–19. DOI of the original paper. The bounded rationality, satisficing and scissors metaphor of Chapter 11 follow, respectively, his decision research and his later synthesis, and do not pass off every sentence of the scissors metaphor as the words of the prize lecture. ↩︎
Falk Lieder & Thomas L. Griffiths (2020). “Resource-rational Analysis: Understanding Human Cognition as the Optimal Use of Limited Computational Resources.” Behavioral and Brain Sciences, 43, e1. PDF of the paper and commentaries. Chapters 11 and 15 use it to describe a line of research that brings cognitive cost into the analysis. The forty-minute study arrangement in the main text is a hypothetical example for illustration, not an intervention this research has tested; nor does the resource-rational framework guarantee that people in fact always allocate their thinking optimally. ↩︎
Gary Klein (1998). Sources of Power: How People Make Decisions. MIT Press, the interview narratives on intuition and fireground commanders. Publisher’s record. Chapter 12 keeps the incidents the researcher recorded and his questioning after the event, kept apart from accident investigation and real-time measurement of the mind; for the conditions of professional reliability see also Kahneman and Klein (2009). ↩︎
Wen Li, Isabel Moallem, Ken A. Paller & Jay A. Gottfried (2007). “Subliminal Smells Can Guide Social Preferences.” Psychological Science, 18(12), 1044–1049. PDF of the paper provided by the authors. Chapter 12 is limited to the effect of odours not consciously detected on specific evaluations, and does not use it to prove that intuition is generally accurate. ↩︎
Michael K. McBeath, Dennis M. Shaffer & Mary K. Kaiser (1995). “How Baseball Outfielders Determine Where to Run to Catch Fly Balls.” Science, 268, 569–573. DOI of the original paper, the authors’ public copy. Chapter 12 discusses continuous visual feedback; the study’s model has its conditions and cannot be turned into “keep the angle of elevation constant and the catch is guaranteed”. Carrying a bowl of water is an everyday example used to illustrate continuous adjustment, and it is not concluded from this that the two activities share the same neural mechanism. ↩︎
Daniel Kahneman & Gary Klein (2009). “Conditions for Intuitive Expertise: A Failure to Disagree.” American Psychologist, 64(6), 515–526. DOI of the paper. Chapter 12 adopts the conditions of environmental regularity and learning feedback, and distinguishes subjective confidence from the reliability of judgement. ↩︎
John R. Anderson (1982). “Acquisition of Cognitive Skill.” Psychological Review, 89(4), 369–406. DOI of the original paper. Chapter 12 uses it for explicit knowledge, practice and procedural ability; it does not assign every emotional change, tacit learning or intuition to the same compilation mechanism. ↩︎
Baruch Fischhoff & Ruth Beyth (1975). “I Knew It Would Happen: Remembered Probabilities of Once-Future Things.” Organizational Behavior and Human Performance, 13, 1–16. PDF of the original paper. Chapter 12 distinguishes the drift of memory that a record of predictions can prevent from the gaps in choice, sampling and counterfactuals that a record cannot fill. ↩︎
J. D. Cole & E. M. Sedgwick (1992). “The Perceptions of Force and of Movement in a Man without Large Myelinated Sensory Afferents below the Neck.” The Journal of Physiology, 449, 503–515. Abstract of the study. Chapter 13 is limited to this case’s loss of sensation, visual feedback and motor performance, and does not extend the case’s results into a condition shared by all sensory disorders. ↩︎
National Heart, Lung, and Blood Institute. “How Blood Clots.” Official explanation. Chapter 13 uses it for the basic distinction between platelets, clotting proteins and the formation of a clot. ↩︎
University of Hawaiʻi. “General Senses and Spinal Cord.” Anatomy and Physiology. Open textbook. The concepts of the reflex arc, the spinal cord and descending motor regulation that Chapter 13 needs can be found in the Motor Pathways and Reflexes sections of that chapter. ↩︎
National Institute of Neurological Disorders and Stroke (2012). Brain Basics: Know Your Brain. PDF of the original NINDS booklet, preserved by UTHealth. Chapter 13 takes only the basic anatomical account of different brain regions taking part and working together, and does not treat the booklet’s simplified introduction to their division of labour as a one-to-one system architecture. ↩︎
Stuart B. Mazzone et al. (2011). “Investigation of the Neural Control of Cough and Cough Suppression in Humans Using Functional Brain Imaging.” Journal of Neuroscience, 31(8), 2948–2958. Study record and abstract. Chapter 13 cites the difference in activity between coughing and its suppression, and keeps the distinction between imaging correlations and a complete causal explanation. ↩︎
Alyn H. Morice et al. (2007). “Opiate Therapy in Chronic Cough.” American Journal of Respiratory and Critical Care Medicine, 175(4), 312–315. Study record and abstract. Chapter 13 uses the finding that the symptom ratings and the citric acid challenge test did not show the same change, to make the point that different measurements cannot be merged into a single threshold narrative. ↩︎
Consumer Financial Protection Bureau (2016), enforcement record for Wells Fargo Bank, N.A.. Chapter 14 draws on the regulatory record to discuss unauthorised accounts, sales incentives and the cost to customers, and does not invent month-by-month changes, the inner lives of individual employees or the effects of every reform since. ↩︎
Michael G. Vann (2003). “Of Rats, Rice, and Race: The Great Hanoi Rat Massacre, an Episode in French Colonial History.” French Colonial History, 4, 191–203. DOI of the study. Chapter 14 retells the episode from this study of the historical archives, and does not add the names, dialogue or personal motives of the rat-catchers, or invented scenes of administrative decision-making. ↩︎
Charles A. E. Goodhart (1975). “Problems of Monetary Management: The U.K. Experience”; Donald T. Campbell (1976). Assessing the Impact of Planned Social Change, original report; Marilyn Strathern (1997). “‘Improving Ratings’: Audit in the British University System.” European Review, 5(3), 305–321, publisher’s page for the paper. Chapter 14 keeps the three authors and the background to their problems apart; the main text paraphrases rather than quotes. The wording now common, “when a measure becomes a target”, is closely related to Strathern’s text of 1997; the contexts and phrasing of Goodhart, Campbell and Strathern each differ, and they should not be treated as different signatures on the same sentence. ↩︎
Ellen J. Weber, Suzanne Mason, Adrian Carter & Rachel L. Hew (2011). “Emptying the Corridors of Shame: Organizational Lessons from England’s 4-Hour Emergency Throughput Target.” Annals of Emergency Medicine, 57(2), 79–88.e1. Study record, DOI of the original paper. The study interviewed emergency-department leaders at nine hospitals between June and August 2008 and was published in 2011. It is an interview study of organisational experience; the changes reported by the interviewees must be distinguished from research that measures clinical outcomes directly. ↩︎
American Psychiatric Association, DSM-5 explanation of autism spectrum disorder, care of LGBTQ patients and the historical background; U.S. Department of Veterans Affairs, PTSD History and Overview. Chapter 14 discusses the history of revisions and the social work that classifications do; it does not offer individual diagnosis, nor does it equate a classification with the benefits paid in each place. ↩︎
Ian Hacking (2007). “Kinds of People: Moving Targets.” Proceedings of the British Academy, 151, 285–318. The author’s lecture paper. Chapter 14 adopts the interaction between a classification and those classified, and does not infer from looping effects that illness or suffering is fictitious. ↩︎
Edwin Hutchins (1995). “How a Cockpit Remembers Its Speeds.” Cognitive Science, 19(3), 265–288. PDF of the original paper. Chapters 14 and 15 adopt its analysis of speed markers and the distribution of cognitive work; it is not used to claim that a group necessarily has a single consciousness. ↩︎
Andy Clark & David J. Chalmers (1998). “The Extended Mind.” Analysis, 58(1), 7–19. Original text provided by the authors. Chapters 14 and 15 list it as a philosophical claim about the boundary of the mind, kept apart from the more limited analysis of external dependence. ↩︎
U.S. Preventive Services Task Force (2009). “Screening for Breast Cancer: Recommendation Statement.” Annals of Internal Medicine, 151, 716–726. Official text of the recommendation from that year. Chapter 15 treats this as a historical case, not as the current screening recommendation. The main text also states explicitly that it uses hypothetical numbers to explain the denominator, and does not treat them as medical risk estimates. ↩︎
Richard Rudner (1953). “The Scientist Qua Scientist Makes Value Judgments.” Philosophy of Science, 20(1), 1–6, publisher’s page for the paper; Heather Douglas (2000). “Inductive Risk and Values in Science.” Philosophy of Science, 67(4), 559–579, DOI of the original paper. Chapter 15 distinguishes the role of values in bearing the risk of error from rewriting facts at will according to values; this is an argued position, not an uncontested definition. ↩︎

