Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Whether character exists

Virtue ethics rests on an empirical bet: that people have stable character traits which explain what they do, and that these traits carry across situations.

If that is false, the theory is in trouble in a way its rivals are not. Consequentialism and deontology can survive the discovery that human character is thin, because they tell you what to do rather than what to be. A theory whose central concept is a settled disposition needs settled dispositions to exist. This lesson is the empirical challenge, and it is unusual in a philosophy course: the evidence is experimental, the numbers matter, and both the challenge and the reply have been damaged by problems in the underlying research that only came to light recently.

The bet, stated precisely

The claim under test is not that people differ. Obviously they do. It is what John Doris calls globalism: that character traits are consistent across a wide range of situations, that they are stable over time, and that they are integrated, so that possessing one virtue predicts possessing others.

Cross-situational consistency is the crucial part, because it is what makes a trait explanatory. If honesty is a trait, then a person who does not cheat on a test should also not lie about their expenses and not steal from a till, and knowing that someone is honest should let you predict what they will do somewhere you have not seen them. Aristotle's account requires exactly this: a virtue is a firm and unchanging state, and its possessor acts well reliably rather than occasionally.

What the personality data says

The first serious test predates the philosophical argument by seventy years. Hugh Hartshorne and Mark May ran the Character Education Inquiry between 1928 and 1930, testing several thousand schoolchildren for honesty across many separate opportunities: copying answers, lying about it afterwards, taking money from a puzzle box, falsifying a self-scored test. If honesty were a trait, performance on one test should predict performance on the others. The correlations they found were small, and their conclusion was that honesty is specific to situations rather than general to persons.

Walter Mischel gathered the accumulated evidence in Personality and Assessment in 1968 and produced the number that has organised the field since. Correlations between a personality measure and behaviour in a single situation rarely exceed about 0.30, a figure often called the personality coefficient.

Work out what that means, because the bare number is often quoted without its interpretation. A correlation of r=0.30 accounts for r2=0.09 of the variance in behaviour, so about 9 percent. Ninety-one percent of the variation in whether someone helps, cheats or obeys is left unexplained by the trait measure. Against that, well-designed situational manipulations routinely move behaviour by twenty or thirty percentage points. On this evidence, if you want to predict what someone will do, you should ask about the circumstances rather than about the person.

The experiments

Three classic studies gave the argument its force, and their numbers are worth having exactly.

Stanley Milgram, in 1963, told subjects they were administering electric shocks to a learner in a memory experiment, with a switchboard rising to 450 volts labelled "danger: severe shock" and then "XXX". The learner, an actor, protested, screamed, complained of a heart condition and eventually fell silent. An experimenter in a lab coat delivered a fixed sequence of prods. In the baseline condition, 26 of 40 subjects, 65 percent, continued to 450 volts. Psychiatrists surveyed beforehand had predicted that about one in a thousand would.

John Darley and Daniel Batson, in 1973, sent Princeton seminary students across campus to give a talk. Half were assigned the parable of the Good Samaritan as their topic. On the way, each passed a man slumped in a doorway, coughing and groaning. What predicted helping was not the topic, which had no significant effect, but how much of a hurry they had been put in: 63 percent of the low-hurry group stopped, 45 percent of the intermediate group, and 10 percent of the high-hurry group. Some students in the high-hurry condition literally stepped over the victim on their way to give a talk about stopping to help a stranger.

Bibb Latané and John Darley, in 1968, seated subjects in a room that began filling with smoke. Alone, 75 percent reported it. With two passive confederates present, 10 percent did. Nothing about the individuals changed; the presence of two people doing nothing suppressed the response almost completely.

The pattern across all three is the same. Trivial and morally irrelevant features of the situation, a man in a coat, three minutes of schedule, two strangers not reacting, move behaviour by amounts no measured trait comes close to matching.

Example. Compute what proportion of behavioural variance a personality coefficient of 0.30 explains, and say why this does not by itself show that traits are unimportant.

The proportion is r2=0.302=0.09, or 9 percent. Two cautions apply before drawing the conclusion. First, 9 percent of variance in a single act is not nothing: an effect of that size, applied across a lifetime of choices, produces very different lives, and the same magnitude in medicine would be a treatment worth having. Second, the comparison with situational effects is not like for like. Situational manipulations are measured as differences between group averages, and trait effects as correlations at the individual level, and a large shift in a group average is compatible with the ordering of individuals within the group being entirely stable. Milgram's own numbers show this: the situation moved compliance far above what anyone predicted, and 14 of 40 subjects, 35 percent, still refused. Something distinguished them, and the experiment was not designed to find out what.

Now you. Darley and Batson found that the assigned topic had no effect on helping. What does this show, and what does it not?

Answer

It shows that having the content of a moral obligation vividly in mind, minutes before the occasion to act on it, does not reliably produce the act, which is a genuinely deflating result for any view on which moral knowledge is the main determinant of moral behaviour. It does not show that the students had no relevant traits, for two reasons. The measure is a single act in a single situation, which is exactly the design Mischel's ceiling applies to, and the manipulation that worked was one that changed how the situation was perceived: a hurried person may not have registered the man as needing help at all. That second point matters for the virtue tradition specifically, because its claim is precisely that virtue is a matter of perception, of noticing what a situation contains, and a study showing that attention is easily disrupted is describing the mechanism the theory says is central rather than refuting the theory.

The philosophical argument

Gilbert Harman drew the strong conclusion in 1999. Ordinary attributions of character traits are systematically mistaken, in the way that attributions of witchcraft were: we explain behaviour by invented dispositions because we commit the fundamental attribution error, over-weighting the person and under-weighting the situation. If there are no character traits, virtue ethics has no subject matter and should be abandoned.

John Doris, in Lack of Character in 2002, argued for something more careful and more damaging. The evidence tells against globalist traits, the broad cross-situational ones the tradition needs. It is compatible with local traits, narrow and highly situation-indexed dispositions: not courage but something like sailing-in-rough-weather-with-friends courage, which may be perfectly real and perfectly stable while telling you nothing about how the same person behaves in a burning building. Local traits exist, they are numerous and fragmented, and they do not add up to the unified character that eudaimonist ethics is built on.

Doris also draws a practical moral that is worth more than the theoretical one. If situations are this powerful, the reliable route to good behaviour is not to build character and trust it, but to attend to circumstances: avoid situations you know to be corrupting, arrange your commitments so that you are not hurried past the person who needs help, and design institutions on the assumption that ordinary people will do what the setting invites. Anyone who thinks their integrity will hold under a Milgram-strength situation has misread the evidence about people in general and has no special evidence about themselves.

The replies

Three replies are strong enough to keep the debate open, and one of them is arithmetical.

The aggregation reply, made by Seymour Epstein in 1979, points out that a correlation against a single act is the wrong measure. Any single behaviour is noisy, and the standard psychometric correction applies: aggregating over k occasions raises the observed correlation to kr/(1+(k-1)r). Starting from r=0.30, aggregating over 5 occasions gives 0.68, and over 10 occasions gives 0.81. The ceiling is largely an artefact of judging traits by one-shot criteria, and the same correction is applied without controversy to test items and to medical measurements. Since virtue is a claim about how someone acts across a life rather than on one afternoon, the aggregated figure is the relevant one.

The rarity reply, pressed by Rachana Kamtekar in 2004, notes that Aristotle never claimed most people have virtue. He claims the opposite repeatedly: full virtue is rare, requires long habituation and good fortune in one's upbringing, and most people are at best continent. Experiments on unselected undergraduates therefore test a population the theory predicts will mostly fail, and finding that they mostly fail is not a refutation. The theory would be refuted by evidence that nobody can achieve stable good conduct, and there is no such evidence.

Example. Distinguish what Harman and Doris each conclude, and say what evidence would tell against each.

Harman concludes that character traits do not exist and that attributing them is a systematic error like attributing witchcraft. Doris concludes that broad cross-situational traits are unsupported while narrow situation-indexed ones are real and numerous. Evidence against Harman is easy to specify and already exists: any demonstration that a measured disposition predicts aggregated conduct refutes him, and the aggregation results do exactly that, which is why almost nobody holds his position now. Evidence against Doris is harder, and this is what makes his version the serious one. It would have to show consistency across genuinely dissimilar situations, not merely repeated occasions of the same kind, since repeated occasions are what aggregation aggregates. Longitudinal work showing that a trait measured in one domain predicts conduct in an unrelated one, above the level Mischel's ceiling allows, would do it. Noticing that the two positions differ this much in what would refute them is the point of the exercise: they are usually cited together, and only one of them is still standing.

Now you. Someone says the situationist results show that people are not really responsible for what they do, since the situation caused it. Assess this.

Answer

It overreaches in a way the free will course diagnosed. That a situation raises the proportion of people who do something from 5 percent to 65 percent does not show that any individual was compelled: in Milgram's baseline, 35 percent refused under exactly the same pressure, so the situation made the act much more likely and left it possible to resist. What the results do support is a narrower and still important claim about excuses, which is that a person acting under strong situational pressure is less culpable than one acting freely, and that this is a matter of degree rather than an all-or-nothing exemption, which is what ordinary moral practice already holds about duress. They also support a claim about what to do rather than about who to blame: since we know situations are this powerful, arranging not to be in them is itself something a person can be held responsible for, and so is designing institutions that do not manufacture them.

The profile reply comes from Mischel himself, who did not conclude that persons do not matter. With Yuichi Shoda in 1995 he proposed that consistency lies in if-then signatures: a person is reliably more aggressive than others when criticised by a peer and reliably less aggressive when criticised by an adult, and that pattern is stable over years even though their average aggression predicts little. On this account character is real and is a structure rather than a quantity, which is closer to Aristotle's talk of acting rightly towards the right people at the right time than the crude trait model either side was arguing about.

The evidence has its own problems

An honest account has to add that some of the material on both sides of this argument has not held up, and the recent history is a caution against citing any of it too confidently.

The Stanford Prison Experiment of 1971, for decades the most cited demonstration of situational power, has collapsed. Thibault Le Texier's examination of Philip Zimbardo's own archives, published in American Psychologist in 2019, showed that the guards were briefed on the behaviour expected of them rather than inventing it, that instructions to be tough were given and repeated, that the most notorious guard later said he was consciously playing a role, and that the study had a conclusion before it had data. It should no longer be cited as evidence of anything.

Milgram's work survives in outline and is more complicated than its summary. Gina Perry's examination of the Yale archives showed that experimenters departed from the scripted prods, that some subjects doubted the shocks were real, and that the frequently quoted 65 percent comes from one of more than twenty conditions whose results ranged from near-total compliance to near-total refusal. The variation is itself informative: obedience fell sharply when the experimenter gave orders by telephone, when the learner was in the same room, and when other subjects refused first. Jerry Burger's partial replication in 2009, stopped at 150 volts for ethical reasons, found rates close to Milgram's, so the basic phenomenon is real.

Isen and Levin's 1972 finding, that 14 of 16 people who found a dime in a phone booth helped a stranger against 1 of 24 who did not, is the single most quoted situationist result and has a poor replication record. It should be treated as an illustration rather than as evidence.

The general lesson is one this course applies elsewhere: when a philosophical position leans on an empirical literature, it inherits that literature's reliability, and social psychology's reliability has been revised downwards over the last fifteen years.

Example. Starting from a personality coefficient of 0.30 for a single occasion, compute the correlation expected after aggregating over ten occasions, and say what the result licenses.

Apply the aggregation formula with k=10 and r=0.30: the numerator is 10×0.30=3.0, the denominator is 1+9×0.30=3.7, and the ratio is 3.0/3.7=0.81. That is a strong relationship. What it licenses is the claim that traits predict aggregated conduct well, which is the claim virtue ethics actually makes, since nobody has ever said that a courageous person acts courageously on every single occasion. What it does not license is any confidence about a particular occasion, and this is the practical residue of the situationist argument that survives every reply: you cannot rely on anyone's character, including your own, to withstand a strong situation on a given day. Both conclusions are true together, and most of the heat in the debate came from treating them as competitors.

Now you. A company wants to reduce fraud among its staff. What does the situationist evidence recommend, and what does the virtue tradition add?

Answer

The situationist recommendations are all about circumstances, and they are the ones that work: remove opportunities, separate the person who authorises from the person who pays, require signatures at the moment of assertion rather than after, make the behaviour of others visible so that nobody can assume a norm of quiet tolerance, and avoid targets that put people under the kind of pressure the hurried seminarians were under. Hiring for integrity, by contrast, has weak evidence behind it, exactly as the personality coefficient predicts. The virtue tradition adds two things it can consistently claim. Habituation is real, so a workplace in which small honesty is practised and expected shapes what people become over years rather than merely constraining what they do today. And what people notice is trainable: most workplace fraud begins with something the person did not see as fraud at all, and the tradition's claim that moral failure is usually a failure of perception rather than of will is, in this domain, well supported.

Where this leaves the theory

Virtue ethics has been made more modest by this literature and has not been dislodged.

What it has to give up is the strong globalist picture in which knowing that someone is honest tells you what they will do anywhere, and any suggestion that a well-formed character is armour against circumstances. What it keeps is the claim that dispositions are real and predict aggregated conduct, that they are formed by habituation, that virtue is rare and hard, and that moral perception is central. Each of those is either supported by the evidence or untouched by it, and the last one is arguably strengthened, since the mechanism by which situations work is that they change what people notice.

The practical upshot is a hybrid that most careful writers in the area now hold: cultivate character, and do not trust it. Build the dispositions, and also build the institutions and habits that make the dispositions less load-bearing.

All three normative families are now on the table, with their strengths and their unpaid debts. Every one of them, though, has been quietly assuming an answer to a prior question that none of them has argued for: whose interests belong in the calculation at all. That question is next, and the answer turns out to be worth billions of lives.