AP Psychology · Science

AP Psychology

The redesigned AP Psychology course, taught the way the exam actually tests it: five units of concepts, each one anchored to the research that produced it. Every lesson defines the vocabulary, walks through a real study (method, participants, findings, one number), and then makes you apply it to a scenario. Each unit ends with a ten-question review in the exam's three item styles, and the two free-response sections let you practice the Article Analysis Question and the Evidence-Based Question against the College Board rubrics with a full-credit model to compare against.

2H 40M EXAM 75 MCQ 2 FRQ 51 LESSONS 50 REVIEW QUESTIONS 10 FRQ PROMPTS NO PREREQ

Course overview

What this course covers, and how the exam weights it.

AP Psychology is organized into five units, and the College Board weights each one at 15–25% of the exam, no unit is safe to skip. The exam is 2 hours 40 minutes and digital: Section I is 75 multiple-choice questions in 90 minutes (66.7% of the score), and Section II is two free-response questions in 70 minutes (33.3%): the Article Analysis Question, worth 7 points in 25 minutes, and the Evidence-Based Question, worth 7 points in 45 minutes. Because both free-response questions hand you research to read, research methods and statistics are not a separate unit here; they are woven through all five as their own lessons, so you meet operational definitions, correlation, effect size, sampling, p-values and ethics in the units where they matter. Every unit below ends with a ten-question review that mixes the three multiple-choice styles the exam uses: concept application, data analysis, and scientific investigation. The two writing sections at the bottom of the outline hold five AAQ prompts and five EBQ prompts, each with the research summaries, a timed writing space that scores against the official rubric, and a 7/7 model response.

  • U1Biological Bases of Behavior15–25%
  • U2Cognition15–25%
  • U3Development and Learning15–25%
  • U4Social Psychology and Personality15–25%
  • U5Mental and Physical Health15–25%

All five units are open, 51 lessons in all. Every lesson pairs a short explanation with worked examples and a problem to try yourself, the same problem types that show up on the exam. Each unit closes with a short video walk-through and a ten-problem practice set with hidden answers.

Free preview: open any 5 lessons, or watch one unit video, without an account. The counter on the left keeps track.

Lesson 1.1 · Unit 1 · Research methods

How psychologists ask questions: hypotheses and designs

Everyone has theories about people: that opposites attract, that music helps you study. Psychology's advantage is procedure: it states a claim precisely enough that some result could prove it wrong, then collects that result. The design then fixes what you may conclude; experiments alone establish cause.

Key terms
  • Theory and hypothesis: a theory explains and predicts; a hypothesis is a single falsifiable prediction from it.
  • Independent and dependent variable: what the experimenter manipulates, and the outcome measured.
  • Control group and placebo: an untreated or placebo group supplies the baseline.
  • Random assignment vs. random selection: assignment sorts participants into conditions by chance, permitting a causal claim; selection draws the sample, permitting generalization.
  • Confound and blinding: a confounding variable changes along with the independent variable; double-blind procedures hide condition from participants and researchers.
  • Non-experimental designs: correlational studies, naturalistic observation, case studies, surveys, and meta-analyses describe relationships rather than manipulate them.
Study spotlight

Study: Dunn, Aknin & Norton (2008)

Researchers approached students on a Canadian university campus, rated their happiness, and handed each one an envelope holding either $5 or $20. Each was randomly assigned to spend it before 5 p.m. on themselves or on someone else. That evening the researchers contacted them and rated happiness again. Students who spent on themselves finished no happier than they began; students who spent on someone else reported higher happiness. The amount changed nothing: $20 bought no more happiness than $5.

Reading the research: An experiment: spending target and amount were manipulated and participants randomly assigned, which is what licenses causal wording. The dependent variable was operationalized as a self-reported happiness rating that evening. The limitation is the sample: fewer than fifty undergraduates at one university.

Apply it

Marcus insists studying with music helps him learn. Turn the claim into a study.

Make it measurable: students who study a word list with music will recall more words than students who study it in silence. The independent variable is the presence of music, the dependent variable is words recalled twenty minutes later, silence is the control. Assign randomly, if students pick, preference becomes a confound.

Exam tip: concept application items ask which element is the dependent variable. Read for what was measured.

Write it

Identify the research method used in this study and explain one design feature that allows a causal conclusion.

Show a model response

The researchers used an experiment. Participants were randomly assigned to spend on themselves or on someone else, so the groups should not have differed beforehand in generosity or mood. Because spending condition was the only thing manipulated, the difference in evening happiness can be attributed to it.

Why it earns the point: it names the method, then cites random assignment instead of retelling the procedure.

Lesson 1.2 · Unit 1 · Research methods

Operational definitions and measurement

Psychology studies things nobody can hand you: happiness, intelligence, aggression, grit. These are constructs: ideas with no ruler attached. The fix is procedural: say exactly what you will count. A study is only as good as the sentence that turns its construct into a measurement.

Key terms
  • Construct: a variable inferred rather than directly observed, such as anxiety or motivation.
  • Operational definition: the exact procedure used to measure a variable, stated so someone else could repeat it.
  • Reliability: consistency: test-retest reliability across occasions, inter-rater reliability across coders.
  • Validity: whether the measure captures the construct (construct validity) and forecasts what it should (predictive validity).
  • Self-report bias: social desirability pushes people toward flattering answers; a response set answers in a fixed pattern.
  • Population vs. sample: everyone the conclusion is about, versus who was actually studied.
Study spotlight

Study: Mehl, Vazire, Holleran & Clark (2010)

Seventy-nine undergraduates wore an electronically activated recorder for four days. It sampled 30 seconds of ambient sound every 12.5 minutes, so participants could not manage what it caught. Coders classified each clip as small talk (weather, what's for lunch), or as substantive conversation about something of consequence. Well-being was rated separately, by the participants and by people who knew them. The happiest participants had roughly a third as much small talk as the unhappiest, and about twice as many substantive conversations.

Reading the research: Correlational, not experimental: nobody was assigned to talk more deeply, so the direction is open. The strength is the operational definition, a coded clip rather than a memory of one; the limitation is a sample of undergraduates at one university.

Apply it

Define "screen time" and "sleep quality" so another researcher could reproduce your numbers.

Screen time: total minutes of foreground app use, read off the phone's own usage screen, summed midnight to midnight for seven days. Sleep quality: the mean answer, on a 1–5 scale, to "how rested did you feel on waking?" entered on waking across those seven days. Each names the instrument, the window, and the scale. "Slept well" is not an operational definition, because two researchers would count it differently.

Exam tip: AAQ part B asks you to state an operational definition. Give the procedure, not the concept.

Write it

State how well-being was operationalized in this study, and name one threat to the validity of that measure.

Show a model response

Well-being was operationalized as ratings of happiness and life satisfaction given by the participants themselves and by acquaintances who knew them. One threat is social desirability: participants may report being happier than they feel, which would inflate scores and shrink the real gap between groups.

Why it earns the point: it reports the procedure, not a definition of happiness, then names a specific bias.

Lesson 1.3 · Unit 1 · Research methods

Research ethics and the APA guidelines

Every rule in this lesson exists because a study once harmed someone. Ethics is not decoration on a finished design; it is a gate the design must pass first: a committee weighs what a study could teach against what it costs the people in it, and can say no.

Key terms
  • IRB and IACUC: boards that review research with human participants and with animals before it may run.
  • Informed consent and assent: adults agree after learning the risks; a minor gives assent while a guardian consents.
  • Deception: allowed only when the study cannot work otherwise and participants are debriefed.
  • Protection from harm: physical and psychological risk must be minimized.
  • Confidentiality and anonymity: confidential data are linked to a name but guarded; anonymous data record none.
  • Debriefing and right to withdraw: participants may stop without penalty, and learn the true purpose afterward.
Study spotlight

Study: Milgram (1963)

Forty men from the New Haven area, recruited by newspaper advertisement, were told the study concerned learning and memory. Each took the role of teacher at a generator with switches labeled from 15 to 450 volts. When the learner, a confederate, erred, an experimenter in a lab coat directed the teacher to raise the shock level, prodding him onward whenever he objected. No shocks were actually delivered. Twenty-six of the 40 men, 65%, continued to the maximum level.

Reading the research: A laboratory experiment with no control group: every participant met the same escalating instructions, and the dependent variable was the highest shock level delivered. The limitation is ethical as much as methodological: participants were deceived and visibly distressed, which helped produce today's IRB review and debriefing rules.

Apply it

A student proposes secretly recording cafeteria conversations to study gossip. What breaks, and what would an IRB demand?

Recording people who never agreed to be studied violates informed consent, and keeping voices that identify speakers violates confidentiality. There is no right to withdraw either, since nobody knows they are in a study. An IRB would require consent, transcripts coded so no voice is kept, and a debriefing that lets people delete their data.

Exam tip: AAQ part D asks which ethical guideline the researchers applied. Name it, then point to the procedural step that shows it.

Write it

Identify one ethical guideline Milgram's procedure would violate today, and explain why.

Show a model response

Protection from harm. Participants were pressed to continue while they sweated, trembled, and asked to stop, so the distress went well beyond ordinary life and had not been disclosed. An IRB today would require that risk stated in advance, and a request to stop honored.

Why it earns the point: it names a guideline by its term and ties it to a specific procedural feature.

Lesson 1.4 · Unit 1 · CED topic 1.1

Heredity, environment, and the evolutionary perspective

Nature versus nurture is a badly posed question. Genes never act outside an environment, and environments act on organisms that already have genes. The answerable question is how much of the variation in a trait tracks genetic differences.

That distinction carries history. The eugenics movement read heritability as destiny and used it to justify forced sterilization and exclusion. It was indefensible and statistically wrong, and psychology rejects it.

Key terms
  • Genetic predisposition: an inherited tendency that raises the odds of a trait in certain environments; nature and nurture interact.
  • Heritability: the share of a population's variation in a trait that tracks genetic differences. A group statistic, never a personal one.
  • Twin and adoption studies: compare identical with fraternal twins, and adoptees with both sets of relatives, separating shared genes from shared homes.
  • Epigenetics: experience switches genes on or off without altering DNA.
  • Natural selection: traits aiding survival and reproduction spread; the evolutionary perspective explains behavior by its ancestral function.
Study spotlight

Study: Rosenzweig, Bennett & Diamond (1972)

Littermate rats were assigned by chance to one of two housing conditions. Some lived alone in bare cages; others lived in groups in large cages stocked with objects swapped on a regular schedule. After weeks the researchers examined the brains. Rats from the enriched cages had a heavier, thicker cerebral cortex than their impoverished littermates, a difference on the order of a few percent, with greater acetylcholine-related enzyme activity. Chance assignment of littermates rules out genetics as the cause.

Reading the research: An animal experiment. Housing is the independent variable, cortical weight and thickness the dependent variable, and random assignment of littermates rules out heredity. The limitation is generalization: caged rodents are far from a child's classroom.

Apply it

A parent says musical talent is "just genetic," so lessons would be wasted.

Heritability describes variation in a population, not a ceiling on a person. If a musical ability measure is 50% heritable, half the differences among people track genes; it does not mean half of this child's skill is fixed. And where every child gets lessons, remaining differences look more genetic simply because environment stopped varying.

Exam tip: heritability is about variance between people, never about an individual.

Write it

Explain the extent to which this study's findings generalize to humans, citing specific evidence about the participants.

Show a model response

Generalization is limited. The participants were laboratory rats housed either alone in bare cages or in groups with changing objects: conditions far more extreme than the gap between two children's homes. That experience alters brain structure has human support elsewhere, but this result cannot be extended beyond rodents.

Why it earns the point: part E wants participant evidence, so it names the participants and why they limit the claim.

Lesson 1.5 · Unit 1 · CED topic 1.2

The nervous system, the endocrine system, and the neuron

Your body runs two messaging services at once. The nervous system sends fast, targeted electrical signals down wires; the endocrine system floods the bloodstream with hormones that act slowly and everywhere. Fear needs both: a jump in a quarter second, and a jittery half hour after.

Key terms
  • Central and peripheral nervous system: brain and spinal cord; the nerves outside them.
  • Somatic and autonomic divisions: voluntary muscle control; involuntary organ control, split into sympathetic (arousal) and parasympathetic (calm).
  • Neuron types: sensory (afferent) neurons carry input in, motor (efferent) neurons carry commands out, interneurons link them; a reflex arc skips the brain.
  • Neuron anatomy: dendrites receive, the axon carries, myelin speeds transmission, terminal buttons release neurotransmitter.
  • Glial cells: support cells that insulate, feed, and clean up.
  • Endocrine system: hormone glands: the pituitary directs the others, the adrenals release epinephrine and cortisol.
Study spotlight

Study: Loewi (1921)

Loewi isolated two frog hearts, each beating in its own bath of fluid. He stimulated the vagus nerve of the first heart, and that heart slowed. He then drew off the fluid around the slowed heart and applied it to the second, whose nerve had not been touched, and the second heart slowed too. Something released into the fluid by the stimulated nerve, not the electrical current itself, had carried the message. The substance was later identified as acetylcholine.

Reading the research: A controlled laboratory experiment. Nerve stimulation is the manipulation, heart rate the measure, and the untouched second heart is the comparison that isolates a chemical cause. The limitation is the preparation: two excised frog hearts are a long way from a living brain.

Apply it

Trace your autonomic divisions before, during, and after a class presentation.

Waiting your turn, the sympathetic division fires: the adrenal glands release epinephrine, heart rate and breathing climb, digestion pauses. During the talk that arousal is still running, which is why your hands shake. When you sit down the parasympathetic division takes over, but the epinephrine already in your blood clears slowly, so you stay jittery for minutes after the fear is gone.

Exam tip: items ask which division more than which structure. Sympathetic mobilizes, parasympathetic restores.

Write it

Identify Loewi's research method and state the dependent variable as he operationalized it.

Show a model response

Loewi used an experiment. The dependent variable was operationalized as the beating rate of an isolated frog heart, whether the second heart slowed after receiving fluid from the first. Because that heart was never stimulated, any change in its rate had to come from a chemical in the transferred fluid.

Why it earns the point: it names the method and gives the dependent variable as a measured rate, not as "nerve signals."

Lesson 1.6 · Unit 1 · CED topic 1.3

Neural firing and neurotransmitters

A neuron is not a dimmer switch. It waits at a negative charge, collects incoming signals, and either fires at full strength or not at all. Everything you feel is built from the timing and chemistry of those identical pulses.

Key terms
  • Resting potential and threshold: the neuron sits near −70 mV; excitatory input depolarizes it to roughly −55 mV.
  • Action potential, all-or-none: past threshold the neuron fires at full strength. A stronger stimulus means more frequent firing, not a bigger spike.
  • Refractory period: a brief window after firing when it cannot fire.
  • Synapse and reuptake: neurotransmitter crosses the gap and binds receptors, then is broken down or reabsorbed.
  • Agonist and antagonist: an agonist mimics or boosts a transmitter; an antagonist blocks it.
  • The messengers: acetylcholine (movement, memory), dopamine (reward), serotonin (mood, sleep), GABA (inhibitory), glutamate (excitatory), norepinephrine (alertness), endorphins (pain relief), substance P (pain), oxytocin (bonding).
Study spotlight

Study: Olds & Milner (1954)

Electrodes were implanted in the septal area and nearby forebrain regions of rats, each animal then placed in a box with a lever. Pressing it delivered a brief pulse of electrical stimulation to the animal's own brain. Rats wired to these sites pressed repeatedly and fast, rates in the hundreds to a couple of thousand presses an hour, and kept pressing instead of eating. Rats whose electrodes sat elsewhere showed no such preference, so the effect depended on exactly where in the brain the electrode had been placed.

Reading the research: An animal experiment: electrode placement is the independent variable, lever presses per hour the dependent variable. The limitation is interpretive: it locates a circuit an animal will work for, but says nothing about what the rat experiences.

Apply it

Why does a reuptake-inhibiting antidepressant change signaling at once but mood only weeks later?

The drug blocks the transporter that pulls serotonin back into the sending neuron, so serotonin lingers in the synapse and keeps binding receptors: within hours of the first dose. But mood is not one synapse. Over weeks receptors adjust and downstream circuits reorganize; the clinical change tracks that slower adaptation. Which is why a patient is told not to quit after a week.

Exam tip: answer at the level asked; "more serotonin in the synapse" is neural, "symptoms improve" behavioral.

Write it

Describe what the lever-press rate indicates about the stimulated areas.

Show a model response

The high press rate indicates the stimulation was reinforcing: the rats repeated whatever produced it, the operational signature of reward. The rate does not indicate pleasure: only that the animals worked to repeat the stimulation.

Why it earns the point: part C asks what a statistic indicates, and this reads it as evidence of reinforcement without claiming the rat felt something.

Lesson 1.7 · Unit 1 · CED topic 1.4

Drugs and neural firing

Psychoactive drugs do not invent new experiences. They work by imitating, blocking, or lingering in systems the brain already runs, which is why their effects are predictable from the synapse, and why the brain pushes back against them.

Key terms
  • Psychoactive drug: a substance that crosses the blood-brain barrier and alters neural signaling, and so behavior.
  • Agonist, antagonist, reuptake inhibitor: mimic a transmitter, block its receptor, or slow its removal from the synapse.
  • Tolerance: repeated use changes receptors, so the same dose does less.
  • Withdrawal: when the drug stops, the adapted system is unopposed and produces symptoms roughly opposite to the drug's effects.
  • Substance use disorder: a DSM-5-TR diagnosis naming a pattern of impaired control, social impairment, risky use, and tolerance or withdrawal.
  • Drug classes: depressants slow neural activity, stimulants speed it, opioids mimic endorphins, hallucinogens distort perception.
Study spotlight

Study: Alexander and colleagues (1978–1981)

Rats were housed either alone in standard laboratory cages or together in a large enclosure, called Rat Park, with room to move, nest, and socialize. Over a period of weeks every animal could drink from two bottles: plain water, or water containing morphine, sweetened to mask the bitterness. The isolated rats drank far more of the morphine solution than the colony rats did. The researchers argued that housing conditions, and not the drug's pharmacology alone, shaped how much was consumed.

Reading the research: An experiment: housing is the independent variable, amount of morphine solution consumed the dependent variable. The limitation is replication: samples were small and later attempts mixed, so read it as evidence that environment matters, not as settled.

Apply it

A student says their morning coffee "stopped working," and skipping it gives them a headache.

Caffeine blocks adenosine receptors, the ones that build sleep pressure. The brain compensates by adding receptors, so at the same dose more adenosine still gets through: tolerance, and the reason the usual cup does less. Skip a morning and those extra receptors sit unblocked, adenosine binds freely, and vessels in the head dilate: the withdrawal headache. Withdrawal symptoms tend to be the drug's effects run backwards.

Exam tip: give the mechanism. "Receptors up-regulate" earns the point; "the body gets used to it" does not.

Write it

Explain the extent to which these findings generalize to human addiction, citing evidence about the participants.

Show a model response

Generalization is limited. The participants were laboratory rats, and the manipulation was cage versus colony: an environmental contrast with no exact human counterpart. The design cannot speak to the social, economic, and psychiatric factors behind human substance use disorders, and later replications have been mixed.

Why it earns the point: part E wants participant evidence, so it names the species and the housing conditions before limiting the claim.

Lesson 1.8 · Unit 1 · CED topic 1.5

The brain: structures, plasticity, and how we study it

The brain is not a set of tidy departments, but damage and imaging both show that specific regions do specific jobs. Learn each structure as an answer to two questions: what stops working when it is gone, and what changes when it is used hard?

Key terms
  • Brainstem: medulla (heartbeat, breathing), pons (sleep, arousal), and the reticular activating system, which gates alertness.
  • Cerebellum: movement, balance, procedural learning.
  • Subcortical structures: thalamus routes sensory input, hypothalamus governs drives, hippocampus forms new memories, amygdala tags threat.
  • Cortex: frontal (planning, motor cortex), parietal (somatosensory), temporal (hearing), occipital (vision). Broca's area produces speech, Wernicke's comprehends it.
  • Corpus callosum and split-brain research: the band joining the hemispheres; cutting it to control seizures revealed what each hemisphere does alone.
  • Methods and plasticity: lesions remove tissue, EEG records rhythms with fine timing, fMRI maps blood flow with fine location. Neuroplasticity is reorganization after experience or injury.
Study spotlight

Study: Maguire and colleagues (2000)

Structural MRI scans compared licensed London taxi drivers, who must pass a demanding examination on the city's street layout, with people who did not drive taxis and were matched on age, sex, and handedness. Regional volumes were measured from the scans and compared. The drivers had greater gray-matter volume in the posterior hippocampus and less in the anterior hippocampus. Among the drivers themselves, posterior volume was larger the more years a driver had spent working the city, a positive correlation with experience.

Reading the research: Not a true experiment, nobody was assigned to drive a taxi, so it is quasi-experimental and correlational. The direction is unsettled: navigating may build hippocampus, or people with more volume may be likelier to qualify and stay.

Apply it

A patient has damage to the left frontal lobe near Broca's area. Predict what follows.

Expect non-fluent aphasia: the patient knows what they want to say, but production is slow and effortful, often reduced to content words with grammar dropped. Comprehension stays largely intact, because Wernicke's area sits further back. Since Broca's area neighbors motor cortex, right-side weakness is common. To confirm, a structural MRI locates the lesion.

Exam tip: EEG answers "when," fMRI answers "where." Match the method to the question.

Write it

Explain why these findings do not show that navigation practice causes hippocampal growth.

Show a model response

The study compared people who had already become taxi drivers with people who had not, and nobody was randomly assigned. A difference could reflect a pre-existing trait that made those people likelier to pass the examination at all. The correlation with years fits growth from practice, but self-selection fits it equally well.

Why it earns the point: it names the missing design feature, random assignment, not just that correlation is not causation.

Lesson 1.9 · Unit 1 · CED topic 1.6

Sleep and dreaming

Sleep is not the brain switching off. It is a scheduled sequence of states, each with its own electrical signature and its own work, and the work includes finishing what you learned that day. Cut the sequence short and you lose the part you needed.

Key terms
  • Circadian rhythm: a roughly 24-hour cycle set by light; the suprachiasmatic nucleus reads light and triggers melatonin.
  • NREM stages: N1 is the drift into sleep, N2 brings sleep spindles, N3 is slow-wave sleep.
  • REM sleep: vivid dreaming with fast brain activity and near-paralysed muscles, hence "paradoxical"; REM lengthens toward morning.
  • REM rebound: REM lost to deprivation is partly made up later.
  • Consolidation and dream theory: sleep stabilizes new memories; activation-synthesis says the cortex builds a story from random brainstem activity.
  • Sleep disorders: insomnia, narcolepsy (sudden sleep attacks), sleep apnea (breathing stoppages), somnambulism (walking in N3), REM sleep behavior disorder (the paralysis fails).
Study spotlight

Study: Walker and colleagues (2002); Mednick, Nakayama & Stickgold (2003)

Participants trained on a finger-tapping task, typing a fixed five-key sequence as quickly and accurately as they could. Retested after a day awake, they were barely faster. Retested after a night of sleep, with no extra practice in between, they typed it roughly 20% faster, a gain that appeared only across sleep. A companion line of work on a visual discrimination task found that a nap containing both slow-wave and REM sleep produced improvement comparable to a full night, while shorter naps lacking REM did not.

Reading the research: An experiment: sleep versus an equal interval of wakefulness is the independent variable, change in correct sequences typed the dependent variable. The limitations are small samples of healthy young adults and very narrow tasks.

Apply it

A student plans an all-nighter before a skills test.

Practising until 3 a.m. adds trials but removes the sleep that turns trials into a stable skill, so they arrive with the day's level and none of the overnight gain. Short sleep also cuts REM out of proportion, since REM clusters toward morning, and the debt is collected as REM rebound the next night. Practice, then sleep.

Exam tip: "consolidation" is the term graders look for: sleep stabilizes what practice encoded.

Write it

State the operational definition of learning used in the finger-tapping study.

Show a model response

Learning was operationalized as the change in the number of correct five-key sequences typed in a fixed interval, from the end of training to a later retest. It is a behavioral count, not a rating of how well-learned the skill feels. Because the measure bracketed either sleep or equal waking time, a difference reflects what happened in between.

Why it earns the point: part B wants the procedure behind the number; unit counted, interval, comparison.

Lesson 1.10 · Unit 1 · CED topic 1.7

Sensation: transduction and the senses

Nothing physical gets into your head. Light, air pressure, and molecules stop at a receptor, which converts them into the only currency the nervous system accepts. Everything after is the brain working from a translation.

Key terms
  • Transduction: receptors convert physical energy into neural firing.
  • Thresholds: the absolute threshold is the least stimulus detected half the time; Weber's law makes the difference threshold a constant proportion.
  • Adaptation and signal detection: constant stimulation fades; detection depends on expectation and the cost of a miss, not strength alone.
  • Vision: rods serve dim light, cones color and detail at the fovea. Three cone types give trichromatic theory; opponent-process explains afterimages.
  • Hearing and body senses: the cochlea transduces vibration; place theory explains high pitch, frequency theory low. Conduction loss is mechanical, sensorineural hair-cell damage. Vestibular tracks the head, kinesthesis limbs.
  • Chemical senses and pain: olfaction and gustation respond to molecules; in gate-control theory the spinal cord can block pain.
Study spotlight

Study: Hubel & Wiesel (1962)

Fine microelectrodes recorded from single neurons in the visual cortex of anesthetized cats while bars and edges of light were projected onto a screen in the animal's field of view. Orientation and direction of movement were varied systematically. Individual cells responded selectively: one fired hard to a bar at a particular angle and barely at all once it was rotated, others only to movement in one direction. The researchers called these cells feature detectors: cortical units tuned to elements of a scene.

Reading the research: A controlled animal experiment. The independent variable is the stimulus shown: orientation and direction; the dependent variable is one neuron's firing rate. The limitation is scope: it explains the coding of edges, not whole objects.

Apply it

A tired lifeguard misses a swimmer in trouble; an anxious parent "sees" one who is not.

Detection is a decision, not only a sensation. Fatigue lowers the lifeguard's sensitivity, and an uneventful shift pushes the criterion toward "nothing there," so a real signal is scored as noise: a miss. The parent's criterion sits very low, because a miss feels catastrophic, so splashing clears it: a false alarm.

Exam tip: signal detection items turn on the criterion, not the eyes.

Write it

A student cannot smell their own house and assumes their nose is broken. Explain, and predict what restores it.

Show a model response

This is sensory adaptation, not damage. Receptors stimulated continuously reduce their response, so a constant background odour drops out of awareness while new odours register normally. The prediction: leave for several hours, the receptors recover, and the smell returns on walking back in, as it does for any visitor.

Why it earns the point: it names the concept, ties it to the scenario, and predicts a testable consequence.

Unit 1 review · 10 multiple-choice

Unit 1 review: Biological Bases of Behavior

Ten questions in the three styles the exam uses (concept application, data analysis, and scientific investigation) one drawn from each lesson in the unit; click an option to see why that choice is right or wrong.

Multiple choice

  1. Described study

    A researcher wants to know whether background music helps students learn vocabulary. She posts a sign-up sheet offering two sessions. Students who sign up for the 4 p.m. session study a 30-word list for 20 minutes in a room with instrumental music playing; students who sign up for the 6 p.m. session study the same list for 20 minutes in silence. Both groups take the same recall test immediately afterward. The music group recalls a mean of 19.4 words and the silence group a mean of 22.1 words. The researcher concludes that background music impairs learning.

    Which of the following is the most serious threat to the researcher's causal conclusion?

    An immediate test narrows what the study can say about durable learning, which is a fair limitation. But the delay was identical in both conditions, and a weakness that applies equally to both groups cannot explain why they differed.

    This is self-selection, and it is the confound that matters here. Students who pick a 4 p.m. slot may differ from students who pick 6 p.m. in sleep, course load, or study habits, and every one of those differences travels with the condition. Random assignment is the design feature that would rule them out.

    A lyrics condition would tell you more about which kind of music matters, but it would leave the existing comparison just as uninterpretable. Adding conditions does not repair groups that differed before the manipulation began.

    You cannot judge whether a 2.7-word gap is real by looking at it; that needs the spread of the scores and a significance test. And even a large, clearly real difference would still have self-selection as a rival explanation.

  2. Data: four measures of sleep quality given to the same 150 adults. Test-retest reliability is the correlation between scores taken two weeks apart. The validity column is the correlation between each measure and total sleep time recorded by a wrist actigraph on the same nights.

    MeasureTest-retest rr with actigraph sleep time
    Single item: "I am a good sleeper" (yes/no)0.880.15
    Nightly 1–5 rating of how rested you feel0.610.24
    Nightly diary of minutes asleep0.790.68
    Bed partner's report of how well you slept0.550.34

    Which measure best illustrates that a measure can be highly reliable and still have poor validity?

    The diary is the measure that works: reasonably consistent across two weeks at 0.79, and by far the closest to the actigraph record at 0.68. It is the example of reliability and validity together, not of the two coming apart.

    The partner's report is the weakest entry on both counts, 0.55 across occasions and 0.34 against the actigraph. A measure that is neither consistent nor accurate cannot illustrate a point that requires one high number and one low one.

    This item is the most consistent of the four across two weeks, at 0.88, and the least related to how long the person actually slept, at 0.15. People answer it the same way every time because it reports a stable self-image, which is exactly how a measure can be perfectly reliable and still miss the construct.

    This measure is mediocre on both counts, 0.61 and 0.24. Because its reliability is the second lowest in the table, it cannot be the example of high reliability paired with weak validity.

  3. Described study

    Researchers recruit 48 first-year college students for what the consent form describes as a study of group problem solving. Each participant is seated with three other people who appear to be fellow participants but are working for the researchers and who, on twelve of eighteen trials, give the same obviously wrong answer to a simple perceptual question. The measure is how often the participant gives the group's wrong answer. Participants are told they may stop at any time, and responses are stored without names. Sessions last about twenty minutes.

    Because the procedure uses deception, which additional safeguard is required?

    Deception is permitted only when a study cannot be run without it and participants are told the truth as soon as the session ends. Debriefing restores what an incomplete consent form could not cover, and it is the step this description leaves out.

    Random assignment is a design feature that supports causal claims, not an ethical guideline. Adding a comparison group would strengthen the inference and do nothing at all about the fact that participants were misled.

    The description already says responses are stored without names, so anonymity is in place. It is a genuine guideline, but it is not the one that the use of deception specifically triggers.

    Telling participants beforehand would eliminate the deception rather than justify it, and the behavior under study would vanish with it. The rule is not that deception must be disclosed in advance but that it must be disclosed afterward.

  4. A news story reports that a large twin study estimated the heritability of height in a Dutch sample at about 0.80. A reader concludes that 80% of his own height was produced by his genes and 20% by his diet. Which of the following best identifies the error?

    Heritability is estimated for physical traits as routinely as for psychological ones, and height is one of the standard textbook examples. The trouble with the reader's conclusion is what the statistic refers to, not which kind of trait it describes.

    Heritability is not a twin correlation read off directly. It is calculated from how much more similar identical twins are than fraternal twins, and it describes a population rather than a match rate for any pair of people.

    Genes and environments do interact, but that does not make heritability meaningless. It stays a well-defined statistic about variation within a specific population at a specific time, which is precisely the part the reader drops.

    The statistic is about variance between people: in this Dutch sample most of the differences in height track genetic differences. It does not partition the reader's own height, and in a population where nutrition varied far more, the same trait would show lower heritability.

  5. During an unexpected fire alarm, Amara's heart pounds and her hands go cold within seconds. The alarm is cancelled after two minutes, but ten minutes later she is still shaky and her heart is still fast. Which of the following best explains why her arousal outlasts the alarm?

    The somatic nervous system carries voluntary commands to skeletal muscle; it is not what raises heart rate or constricts the vessels in her hands. Those organs are under autonomic control, which is the whole reason the division exists.

    This is the two-speed design of the stress response. The sympathetic division acts in a fraction of a second through nerves, and it also tells the adrenal glands to release epinephrine into the blood, where it keeps acting on the heart until it is metabolized: minutes after the danger has passed.

    The parasympathetic division works through the vagus nerve and does not wait on a pituitary signal. The pituitary belongs to the slower cortisol pathway, not to the switch that turns arousal off.

    The two divisions are not locked in a tug-of-war at full strength; the parasympathetic division becomes dominant as arousal subsides. Sustained shakiness is explained by how long a hormone stays in the blood, not by simultaneous maximum firing.

  6. Dmitri brushes a hot pan lightly and feels a mild sting; a moment later he grips it and feels intense pain. Which of the following best describes what differs at the level of the individual sensory neurons?

    This is the all-or-none principle. A neuron either reaches threshold and fires at full strength or does not fire at all, so intensity has to be coded some other way, by how often a neuron fires and by how many neurons are recruited.

    This is the most common misconception about action potentials. Once threshold is crossed the spike is the same size every time; a stronger stimulus does not produce a bigger spike, any more than pressing a doorbell harder makes it louder.

    A resting potential sitting above threshold would mean the neuron fired continuously with no input at all, which would make it useless as a signal. Resting potential stays near −70 mV, and it is depolarization from incoming input that carries the cell to roughly −55 mV.

    The refractory period is the brief window after firing during which the neuron cannot fire again, so lengthening it would lower the maximum firing rate rather than raise the signal. Charge does not accumulate across spikes to build a larger one.

  7. Data: described bar graph of mean fluid consumed per day by two groups of laboratory rats.

    The horizontal axis has two groups: rats housed alone in standard cages, and rats housed together in a large social colony. Each group has two bars, one for a sweetened morphine solution and one for plain water. The vertical axis is mean millilitres consumed per day and runs from 0 to 20. For the isolated group the morphine bar reaches 8.4 mL and the water bar 12.0 mL. For the colony group the morphine bar reaches 2.1 mL and the water bar 18.3 mL.

    Which conclusion is best supported by the graph?

    Add the two bars for each group and both totals come to 20.4 mL. Total fluid intake is effectively held constant across the groups, so it cannot explain a four-fold difference in morphine solution.

    The graph says nothing about addictiveness in general; it compares two housing conditions in one species. The colony animals' low intake is a fact about those rats in that environment, not a property of the drug.

    8.4 plus 12.0 and 2.1 plus 18.3 both equal 20.4 mL, so the groups differ in what they drank rather than in how much. 8.4 of 20.4 is about 41% and 2.1 of 20.4 about 10%: the comparison the display actually licenses.

    Two problems at once: the participants were rats, and even inside the study, housing was manipulated in a laboratory rather than across human lives. A causal claim about substance use disorder in people needs evidence this design cannot supply.

  8. After a stroke, a woman speaks in short, effortful bursts, "want... water... cup", and is visibly frustrated by her own speech. She follows spoken instructions accurately and points to the correct object every time she is asked. Damage to which area best accounts for this pattern?

    Damage to Wernicke's area produces the opposite pattern: speech that flows easily but carries little meaning, paired with impaired comprehension. This woman understands everything said to her, which rules that area out.

    Hippocampal damage blocks the formation of new explicit memories; it does not disrupt the production of speech. She is not failing to remember the word: she knows what she wants to say and cannot get it out.

    The cerebellum supports balance, coordinated movement, and procedural learning. Cerebellar damage can slur articulation, but it does not produce the grammar-stripped, telegraphic output described here.

    Non-fluent aphasia (halting, effortful speech with comprehension largely intact and the person aware of the problem) is the signature of damage to Broca's area. Its position beside the motor cortex also explains why right-side weakness so often accompanies it.

  9. Data: one adult's overnight sleep recording, divided into four successive cycles of about 90 minutes. The table gives the minutes of N3 (slow-wave) sleep and of REM sleep in each cycle; the remaining minutes in each cycle were N1 and N2.

    Cycle1234
    N3 sleep (minutes)4530100
    REM sleep (minutes)8183245

    A student whose sleep follows this pattern goes to bed, sleeps through cycles 1 and 2, about three hours, and then gets up to study. Which of the following is best supported by the table?

    The table shows the reverse. N3 falls from 45 minutes in cycle 1 to none at all in cycle 4, so slow-wave sleep is front-loaded and a shortened night keeps most of it.

    Add the columns she skipped. Cycles 3 and 4 hold 32 plus 45, or 77, of the night's 103 REM minutes, close to 75%, while the two cycles she kept contain 75 of the 85 N3 minutes. That asymmetry is why a short night costs REM first, and why REM rebound tends to follow the next night.

    The four cycles in the table are plainly not identical: N3 runs 45, 30, 10, 0 while REM runs 8, 18, 32, 45. Treating cycles as interchangeable is the assumption these data contradict.

    Consolidation is not confined to one stage; overnight gains on motor sequence tasks are associated with sleep containing both slow-wave and REM periods. Keeping N3 while losing three-quarters of REM is a real cost, not a neutral trade.

  10. In the last hour of a long shift, a baggage screener fails to flag a prohibited item that is clearly visible on the scanner image. Her supervisor then announces that inspectors are running test images through the line, and over the next hour she flags several harmless objects as suspicious. Which of the following best explains the change in her performance?

    An absolute threshold is the weakest stimulus a person detects half the time, and it does not drop because a supervisor makes an announcement. Nothing about her eyes changed; what changed is how much evidence she demanded before saying yes.

    Sensory adaptation is a receptor-level decline in response to unchanging stimulation, and the images on her screen change every few seconds. It also cannot explain a rise in false alarms, which is a decision made about ambiguous evidence.

    Signal detection theory separates sensitivity from the criterion. Fatigue and a long uneventful stretch pushed her criterion toward "nothing there," producing a miss; the warning raised the cost of a miss and dragged the criterion down, and the predictable by-product of catching more real threats is flagging more harmless objects.

    Weber's law describes the proportional size of a just-noticeable difference between two stimuli, such as two weights or two tones. It is about discriminating magnitudes, not about the decision rule a person adopts when the incentives change.

Lesson 2.1 · Unit 2 · CED topic 2.1

Perception and attention

Your eyes deliver a flat, upside-down smear of light; you experience solid objects at measurable distances. Everything in between is perception: built partly from the signal and partly from what the brain expects. Attention is the gate on that construction, which is why something can sit in plain view and never be seen.

Key terms
  • Bottom-up and top-down processing: building a percept from features, versus reading it through schemas and expectation; a perceptual set readies you to see one thing.
  • Selective attention: awareness admits one stream; the cocktail party effect is catching your name in a conversation you were not following.
  • Inattentional and change blindness: missing an unexpected object while attending elsewhere; missing something that changed between two views.
  • Gestalt principles: the mind groups before it labels: figure-ground, proximity, similarity, closure.
  • Depth cues: binocular cues need two eyes (retinal disparity, convergence); monocular cues work in a photograph (linear perspective, interposition, relative size, texture gradient).
  • Constancy: size, shape, and color hold steady as the retinal image changes; still frames in sequence are seen as motion.
Study spotlight

Study: Simons & Chabris (1999)

Observers watched a short video of two teams, one in white shirts and one in black, passing basketballs, and were asked to count the passes made by one team. Partway through, a person in a gorilla suit walked into the scene, faced the camera, and walked off, on screen for several seconds. Asked afterwards whether anything unusual had happened, about half the observers missed it entirely, roughly 46% across conditions in the most-cited report. People who simply watched, with no counting task, saw it nearly every time.

Reading the research: An experiment. Video version and counting difficulty were manipulated; the dependent variable was operationalized as a yes or no, whether the observer reported the unexpected event. The limitation is the setting: a brief clip with an absurd intruder is not a windshield.

Apply it

A driver on a hands-free call looks straight at a cyclist and pulls out anyway.

She was paying attention, to the conversation. Selective attention had committed her processing to the call, so the cyclist fell outside the attended stream and never reached awareness. That is inattentional blindness, not an eye that failed. It also explains why hands-free laws disappoint: the hands were never the bottleneck.

Exam tip: concept application items reward the mechanism by name. "Inattentional blindness" earns the point; "she was distracted" does not.

Write it

Identify the research method used, and explain one design feature that supports the conclusion.

Show a model response

The researchers used an experiment. Observers were assigned to different versions of the task and all saw the same unexpected event, so a difference in noticing can be attributed to the attention demand rather than to the event. A condition with no counting task showed the gorilla is easy to see when attention is free.

Why it earns the point: part A wants the method named, and the evidence is a comparison condition, not a retelling of the procedure.

Lesson 2.2 · Unit 2 · CED topic 2.2

Thinking, problem solving, judgment, and decision making

Most thinking is fast, cheap, and good enough. The shortcuts that make it fast are not random errors: they fail in predictable directions, which is exactly why psychologists can study them. Learn each bias as a shortcut doing its job in the wrong situation.

Key terms
  • Concepts, prototypes, schemas: mental categories, their best examples, and the frameworks that organize them. Assimilation fits new input into a schema; accommodation revises the schema.
  • Algorithm and heuristic: a procedure that guarantees a solution but costs time, versus a shortcut that is usually right and occasionally badly wrong.
  • Availability and representativeness: judging likelihood by how easily examples come to mind, or by resemblance to a stereotype while ignoring base rates.
  • Mental set and functional fixedness: reusing the strategy that worked before; seeing only an object's customary use.
  • Framing and anchoring: the wording of equivalent options changes the choice; a first number drags later estimates toward it.
  • Belief biases: confirmation bias hunts for agreeing evidence, belief perseverance survives the loss of that evidence, hindsight bias makes the outcome feel predictable.
Study spotlight

Study: Tversky & Kahneman (1981)

Participants read about an unusual disease expected to kill 600 people and chose between two programs. One group saw the outcomes written as lives saved: one program saves 200 for certain, the other has a one-third chance of saving all 600 and a two-thirds chance of saving nobody. About 72% took the certain option. A second group saw numerically identical outcomes written as lives lost: 400 die for certain, or a one-third chance that nobody dies. Only about 22% took the certain option.

Reading the research: An experiment. The independent variable is the wording of the description, the dependent variable the program chosen. Because the two versions describe the same outcomes, the shift can come only from the frame. The limitation is that these were hypothetical choices with no stakes, made largely by students.

Apply it

A surgeon can say "90 of 100 patients are alive five years later" or "10 of 100 are dead within five years."

Identical numbers, opposite frames. The survival wording presents a gain, and people take a sure gain, so consent rates rise. The mortality wording presents a loss, and people turn risk-seeking about losses, so more patients gamble on waiting. Nothing about the operation changed; the frame did.

Exam tip: framing items always conceal identical numbers. Check that the options really differ before you explain the choice.

Write it

Describe what the difference between 72% and 22% indicates about these participants' decision making.

Show a model response

Both groups faced the same outcomes, so the 50-point gap indicates the description itself drove the choice. Framed as lives saved, the certain program looked far better than the same program framed as lives lost. That indicates preferences here were built at the moment of choice rather than read off stable values.

Why it earns the point: part C asks what a statistic indicates, so it interprets the gap in this study instead of defining "percentage."

Lesson 2.3 · Unit 2 · CED topic 2.3

Introduction to memory: the three-stage model

Memory is not a single container. Information passes through stores with different capacities and different lifespans, and almost everything is dropped at the first two. That is a feature: a system that kept every sound and glance would have nothing left to think with.

Key terms
  • Three-stage (Atkinson-Shiffrin) model: sensory memory holds raw input briefly, short-term memory holds a few items for seconds, long-term memory stores indefinitely.
  • Sensory memory: iconic memory for vision lasts under a second; echoic memory for sound lasts a few seconds.
  • Short-term memory: roughly seven items for about twenty seconds without rehearsal; chunking packs items into larger meaningful units.
  • Working memory: Baddeley's revision: a phonological loop for sound, a visuospatial sketchpad for images, and a central executive that allocates attention.
  • Explicit and implicit memory: consciously retrieved facts and events, versus skills and conditioned responses that show up only in performance.
  • Parallel processing: many streams run at once, so the narrow limit sits on awareness, not on the whole system.
Study spotlight

Study: Sperling (1960)

A grid of twelve letters, three rows of four, was flashed on a screen for a fraction of a second. Asked to report everything they had seen, participants managed about four or five letters, and reported that more had been there but faded while they were writing. Sperling then sounded a high, medium, or low tone immediately after the flash to name a single row. Under this partial-report cue, participants produced roughly three of the four letters from whichever row was called, implying nearly the whole grid had briefly been available.

Reading the research: An experiment in which each participant served in both conditions; the instruction, whole report or partial report, is the independent variable and letters correctly reported the dependent variable. The partial-report advantage is the evidence for iconic memory. Its limitation is also its finding: delay the tone about a second and the advantage is gone.

Apply it

A server takes a six-person order without writing anything down.

Six separate items sits at the ceiling of short-term capacity, so she does not hold six items. She chunks (three vegetarian mains as one group, two identical drinks as one entry, "the usual" for the regular), and rehearses in the phonological loop on the walk to the kitchen. Ask her an hour later and it is gone, because it was never encoded into long-term memory.

Exam tip: capacity items want a number and a mechanism; about seven items, extended by chunking, not by trying harder.

Write it

State the operational definition of memory used in the partial-report condition.

Show a model response

Memory was operationalized as the number of letters correctly reported from the one row named by a tone sounded immediately after the display vanished. It is a count of correct letters from a cued row, not a rating of how much a participant felt they had seen. Scaling that row score up to three rows is what estimates the size of the store.

Why it earns the point: part B wants the measured procedure; what was counted, and under which cue.

Lesson 2.4 · Unit 2 · CED topic 2.4

Encoding memory: depth, spacing, and testing

Studying feels like exposure: more hours in front of the page, more learning. It is not. What ends up stored depends on what you did with the material while it was in front of you, and on when you came back to it. Both of those are choices.

Key terms
  • Effortful and automatic processing: deliberate rehearsal, versus the effortless logging of space, time, and frequency.
  • Levels of processing: shallow structural encoding (how a word looks), intermediate phonemic encoding (how it sounds), deep semantic encoding (what it means). Deeper encoding leaves a more durable trace.
  • Organization and self-reference: mnemonics, the method of loci, and hierarchies impose structure the material lacks; tying material to yourself encodes it better still.
  • Spacing effect: distributed practice beats massed practice, even holding total study time equal. Cramming is massed practice.
  • Testing effect: retrieving material strengthens it more than rereading it does.
  • Serial position effect: primacy (early items rehearsed into long-term memory) and recency (last items still in short-term memory) both beat the middle.
Study spotlight

Study: Craik & Tulving (1975)

Participants saw words one at a time, each preceded by a question. Some asked about appearance: is the word in capital letters? Some asked about sound: does it rhyme with "train"? Some asked about meaning: does it fit this sentence? Participants answered yes or no, not knowing their memory would be tested. A surprise recognition test followed. Words processed for meaning were recognized far more often than words processed for appearance, roughly three times as often in the sharpest comparison, with the rhyme questions in between.

Reading the research: A repeated-measures experiment: depth of the orienting question is the independent variable, recognition of the word the dependent variable, and every participant met all three depths. Because nobody expected a test, the gap cannot be a study strategy. The honest limitation is a confound: deeper questions also take longer to answer.

Apply it

You have a 40-word vocabulary list and a quiz in six days.

Do not reread it. Split the list into four sets of ten and write, for each word, a sentence about your own week: semantic encoding and the self-reference effect in one move. Study each set on two separate days rather than twice in one sitting: the spacing effect. Quiz yourself with the list covered: the testing effect. Give the last session to the middle of the list, which the serial position effect predicts is weakest.

Exam tip: name the mechanism, not the habit. "Spacing effect" and "retrieval practice" earn points; "study more" does not.

Write it

Explain how these results support or refute the claim that repetition alone determines what is remembered.

Show a model response

They refute it. Every word was presented once, so exposure was held roughly constant; only the question asked about it differed. Words judged for meaning were recognized about three times as often as words judged for appearance. If repetition determined storage, the conditions would have produced equal recognition, so depth of processing is what is doing the work.

Why it earns the point: part F wants specific results tied to the concept, so it cites the comparison and says what the comparison rules out.

Lesson 2.5 · Unit 2 · CED topic 2.5

Storing memory: systems, structures, and amnesia

Long-term memory is not one archive. Different kinds of learning are supported by different structures, which is why a single injury can destroy the ability to remember yesterday and leave the ability to acquire a new skill completely intact.

Key terms
  • Long-term potentiation and consolidation: repeated stimulation strengthens the link between two neurons; consolidation then stabilizes the memory across sleep.
  • Hippocampus: registers new explicit memories and passes them to the cortex. It is the gateway, not the warehouse.
  • Cerebellum: supports implicit memory: procedural skills and conditioned responses.
  • Amygdala and flashbulb memory: arousal tags an event for stronger storage. A flashbulb memory is vivid and confidently held, which is not the same as accurate.
  • Explicit and implicit types: episodic memory for events you lived, semantic memory for facts you know, procedural memory for what your hands know.
  • Amnesia: anterograde amnesia blocks new memories; retrograde amnesia erases older ones. Infantile amnesia is the ordinary absence of episodic memories from the first years.
Study spotlight

Study: Scoville & Milner (1957)

To control severe epilepsy, surgeons removed much of the medial temporal lobe on both sides of the brain of a patient known in the research literature as H.M. The seizures eased, and his intelligence, language, and childhood memories were preserved. What he could no longer do was form new conscious memories; he greeted the same researchers as strangers for decades. Across three consecutive days of mirror tracing, drawing a shape while seeing only its reflection, his accuracy steadily improved, yet each day he said he had never attempted the task.

Reading the research: A longitudinal case study with a single participant. Its depth is unmatched, and its limits follow from the same fact: nothing was assigned, and one brain cannot show what is typical. The dissociation it revealed (explicit memory lost, procedural learning intact) is what reorganized the field.

Apply it

A patient with hippocampal damage is asked to learn a colleague's name, learn a simple tune on the piano, recall a childhood holiday, and learn a new fact from the news.

The tune and the holiday survive. Playing a tune is procedural learning, implicit and supported by the cerebellum, so it improves with practice even if the patient denies having practised. The holiday was consolidated years before the damage and sits in the cortex. The name and the news fact are new explicit memories, episodic and semantic, and both need the hippocampus, so both fail. That pattern is anterograde amnesia.

Exam tip: sort tasks by memory system, not by difficulty. "Implicit" is the word that earns the point.

Write it

Explain the extent to which these findings generalize, citing specific evidence about the participant.

Show a model response

Generalization is sharply limited. The evidence comes from one participant whose surgery removed tissue on both sides of the medial temporal lobe: damage no researcher could ethically create. His preserved intelligence and childhood memories make the dissociation convincing, but they do not make him representative; the finding travels only because later patients with similar damage show the same pattern.

Why it earns the point: part E asks for participant evidence, so it names the single case and the specific damage before limiting the claim.

Lesson 2.6 · Unit 2 · CED topic 2.6

Retrieving memory: cues, context, and state

A memory you cannot produce is not necessarily a memory you lost. Most ordinary forgetting is a failure of access: the file is intact and the search term is wrong. That makes retrieval something you can design for instead of hope for.

Key terms
  • Recall, recognition, relearning: producing information unaided, identifying it among options, and the time saved when you learn it a second time.
  • Retrieval cue and priming: any stimulus linked to a memory at encoding can reopen it; priming is an unnoticed cue activating related material.
  • Encoding specificity principle: a cue helps to the degree it was present when the memory was formed.
  • Context-dependent memory: retrieval improves when the external setting matches the setting at encoding.
  • State-dependent and mood-congruent memory: retrieval improves when the internal state matches; a current mood also biases which memories surface.
  • Tip-of-the-tongue phenomenon: partial retrieval, where the first sound arrives without the word: evidence that access failed, not storage.
Study spotlight

Study: Godden & Baddeley (1975)

Members of a university diving club learned lists of unrelated words in one of two places: standing on the beach, or submerged about fifteen feet underwater in scuba gear. Each diver was later tested for free recall either in the same place or in the other one, so every participant met both matching and mismatching combinations. Recall was clearly better when the two settings matched; when they did not, divers produced roughly 30 to 40% fewer words. A later study using recognition rather than free recall found no comparable effect.

Reading the research: A repeated-measures experiment: the match between encoding and retrieval context is the independent variable, words freely recalled the dependent variable, and each diver served in every combination, which controls for individual memory differences. The limitation is the size of the manipulation: a beach and open water differ far more than a bedroom and a classroom do.

Apply it

A student studies in a noisy dorm room and takes the exam in a silent gym.

Expect a modest cost, not a collapse: those settings differ far less than the divers' did. The mechanism is encoding specificity: the cues present while she studied, the music and the desk and the roommate's talking, are absent in the gym and cannot trigger anything. Two fixes follow. Study at least once under test-like conditions, and self-test from a blank page, which builds retrieval routes that travel with her.

Exam tip: context-dependent items change the surroundings, state-dependent items change what is happening inside the person. Read the stem for which one moved.

Write it

Describe what the difference between the matching and mismatching conditions indicates about retrieval.

Show a model response

The divers learned words in both settings, so the gap indicates that where they were tested controlled access rather than storage. Recovering roughly a third fewer words after a change of setting indicates that cues from the learning environment had become part of the trace and needed to be present to reopen it. It does not indicate the words were erased; the same words returned when the setting matched.

Why it earns the point: part C asks what a result indicates, and this reads the difference as evidence about access rather than about storage.

Lesson 2.7 · Unit 2 · CED topic 2.7

Forgetting and memory distortion

Memory does not replay; it rebuilds. Every retrieval assembles a version out of fragments, general knowledge, and whatever you have heard since, and the rebuilt version feels exactly as vivid as an accurate one. That is why confidence is such a poor guide to accuracy.

Key terms
  • Encoding failure: much of what we "forget" never entered memory at all, because attention was elsewhere.
  • Storage decay: Ebbinghaus's forgetting curve: retention drops steeply at first, then flattens out.
  • Interference: proactive interference is old learning disrupting new; retroactive interference is new learning disrupting old.
  • Retrieval failure and motivated forgetting: the cue is missing, or the material is kept out of awareness.
  • Misinformation effect and source amnesia: information met after an event alters what is later reported; source amnesia keeps the content but loses its origin.
  • Constructive memory: memories are rebuilt at retrieval; imagination inflation raises confidence that an imagined event happened.
Study spotlight

Study: Loftus & Palmer (1974)

Forty-five students watched short films of traffic accidents and then answered questions about what they had seen. On the critical question: how fast were the cars going when they hit each other?: the verb varied across groups, among them "smashed into," "collided," and "hit." Mean speed estimates tracked the verb: about 40.8 mph for "smashed into" and 34.0 mph for "hit." In a second experiment, participants returned a week later and were asked whether they had seen broken glass. There was none in the film. About 32% asked with "smashed" said yes, against about 14% asked with "hit."

Reading the research: Two experiments. The verb in the question is the independent variable; the dependent variables are the speed estimate in miles per hour and, a week later, a yes-or-no report of broken glass. The limitation is the material: a film carries none of the stress of a real crash, and the participants were undergraduates.

Apply it

An officer asks a witness, "How fast was the red car going when it ran the light?"

The question smuggles in two claims: that the light was against that driver, and that the car was moving fast. The misinformation effect predicts the witness may absorb both and later report them as things she saw. Ask instead: "What colour was the light when the red car entered the intersection?" For a jury the lesson is not that witnesses lie; a sincere, detailed, confident account can carry details supplied afterwards.

Exam tip: the misinformation effect requires information arriving after the event. If the detail was never noticed in the first place, the answer is encoding failure.

Write it

Explain how these results support or refute the claim that memory works like a video recording.

Show a model response

They refute it. A recording does not change because of how someone asks about it, but changing a single verb moved mean estimates from about 34.0 to about 40.8 mph for the same film, and a week later roughly a third of the "smashed" group reported glass that was never there. That fits reconstructive memory: the wording of the question became part of what was remembered, which playback cannot do.

Why it earns the point: part F wants specific results plus the concept, so it cites both findings and says what they show about reconstruction.

Lesson 2.8 · Unit 2 · CED topic 2.8

Intelligence and achievement: theories and testing

Intelligence is a construct, not an organ, so arguments about it are partly arguments about what a test ought to measure. Two questions run through the topic: is intelligence one thing or many, and does a score describe a person or a moment?

Key terms
  • General intelligence (g): Spearman's claim that scores on different mental tasks correlate, implying one underlying factor. Thurstone argued instead for several primary mental abilities.
  • Multiple and triarchic theories: Gardner proposed relatively independent intelligences; Sternberg proposed analytical, creative, and practical intelligence.
  • Fluid and crystallized intelligence: reasoning through novel problems, which declines with age, versus accumulated knowledge, which does not.
  • Achievement and aptitude tests: what you have learned, versus predicted capacity to learn. Both need standardization, reliability, and validity.
  • The normal curve: scores are scaled to a mean of 100 with a standard deviation of 15, so about 68% of the norm group falls between 85 and 115.
  • Threats to a fair score: stereotype threat depresses performance when a negative stereotype is salient; test bias means a test predicts less accurately for one group; the Flynn effect is the long-run rise in raw scores.
Study spotlight

Study: Mueller & Dweck (1998)

Across six experiments, fifth graders worked a set of reasoning problems, were told they had done well, and were praised in one of two ways: for ability ("you must be smart at these") or for effort ("you must have worked hard"). Offered a next task, most effort-praised children took the challenging set they might learn from, while most ability-praised children took an easy set that would keep them looking successful. After a hard set everyone struggled with, the ability-praised children enjoyed the work less, wanted to stop, and scored lower on a final set than at the start; the effort-praised children improved.

Reading the research: An experiment. One sentence of praise is the independent variable; the dependent variables are the task chosen and the score on the final set. Random assignment to wording is what licenses the causal claim. The limitations are short laboratory sessions and the much smaller effects later school-wide mindset programs have produced.

Apply it

A student scores 115 on a standardized test with a mean of 100 and a standard deviation of 15.

That is exactly one standard deviation above the mean, roughly the 84th percentile, at or above about 84 of every 100 people in the norm group. What it does not mean: that she is 15% smarter than average, that the number is fixed, or that it forecasts any single performance. A teacher who answers "you're so smart" is running the ability-praise condition; "the way you checked question four was careful" is the version that survives a hard test.

Exam tip: data-analysis items love the normal curve. Convert the score to standard deviations first, then read off the percentage.

Write it

Identify the research method used, and state the operational definition of one dependent variable.

Show a model response

The researchers used an experiment, assigning children at random to one of two praise wordings. One dependent variable was operationalized as the child's recorded choice between two task sets: one described as challenging and likely to teach them something, one described as easy and like the problems just finished. It is an observed choice, not a rating of motivation.

Why it earns the point: part A names the method, and part B gives the procedure that produced the measure instead of defining "motivation."

Lesson 2.9 · Unit 2 · Research methods

Descriptive statistics: center, spread, and the normal curve

A set of scores is not a finding until you say something about its shape. Two numbers do most of the work (one for where the scores sit, one for how far apart they are), and picking the wrong one is how honest data end up misleading people.

Key terms
  • Measures of center: the mean follows every score, including extreme ones; the median is the middle value and resists them; the mode is the most frequent score.
  • Measures of spread: the range is highest minus lowest; the standard deviation, the square root of the variance, is roughly how far scores sit from the mean.
  • Shape: a normal distribution is symmetric and bell-shaped; positive skew has a long right tail, negative skew a long left tail; a bimodal distribution has two peaks.
  • Outlier: an extreme score that drags the mean toward it and inflates the range while barely moving the median.
  • Percentile rank: the percentage of scores in the comparison group at or below a given score.
  • z-score: how many standard deviations a score sits from the mean. In a normal distribution about 68% of scores fall within ±1 SD and about 95% within ±2 SD.
Study spotlight

Study: Terman (1921 onward)

California teachers nominated students they judged unusually able; those nominated were then given the Stanford-Binet, and about 1,500 children who scored roughly 135 or above were enrolled. The group's mean score was near 151. Terman and his successors followed them for decades with questionnaires, interviews, and records of schooling, work, and health. As adults the group did well on average, with more education and professional employment than the general population, but the distribution of their achievements was strongly skewed: a few produced most of the notable output.

Reading the research: A longitudinal descriptive study with no control group, so nothing in it is causal. The selection rule is what to name: nomination came before the test, so a child who would have scored above 135 but was never nominated could not enter. The mean of 151 describes a filtered sample, not high-scoring children in general.

Apply it

Two classes both average 78 on the same test. In class A the standard deviation is 4; in class B it is 14.

Same center, different spread. In class A roughly two-thirds scored between 74 and 82, so the mean describes nearly everyone. In class B one standard deviation runs from 64 to 92, so the mean describes almost nobody. If class B is also positively skewed by a few very high scores, report the median: the mean is being pulled by the tail. And "84th percentile" is a position, not a score: 84% of the comparison group scored at or below it.

Exam tip: when a stem says "skewed" or shows an outlier, the median is the defensible choice, and graders want the reason: the mean follows the tail.

Write it

Explain the extent to which these findings generalize, citing specific evidence about the participants.

Show a model response

They generalize poorly. The participants were about 1,500 California schoolchildren who were first nominated by teachers and then scored roughly 135 or above, with a group mean near 151. Nomination filtered the pool before anyone was tested, so quiet or newly arrived children who would have scored highly are missing, and the sample is one state in one era.

Why it earns the point: part E wants participant evidence, so it names the size, the selection rule, and the setting before judging how far the findings travel.

Lesson 2.10 · Unit 2 · Research methods

Correlation: scatterplots, r, and what it cannot tell you

Correlation is the workhorse of psychology, because most of what we want to know cannot be assigned to people. It is also the most misread statistic on the exam. The questions are almost never about computing r; they are about what r does not entitle you to say.

Key terms
  • Correlation coefficient (r): a number from −1 to +1. The sign gives direction, the distance from zero gives strength; r = 0 means no linear relationship.
  • Scatterplot: one dot per case, plotted on both variables. A tight sloping cloud is a strong relationship; a round cloud is none.
  • Directionality problem: a correlation between A and B is equally consistent with B causing A.
  • Third-variable problem: some unmeasured variable may be producing both.
  • Illusory correlation: perceiving a relationship the data do not support, usually because the confirming cases are the memorable ones.
  • Regression toward the mean: extreme scores tend to be followed by less extreme ones, which makes useless interventions look effective.
Study spotlight

Study: Diener & Seligman (2002)

Researchers screened 222 undergraduates on several measures of happiness and compared three groups: the happiest 10%, a middle group, and the least happy 10%. All were assessed on personality, on mood recorded repeatedly over time, on the quality of their relationships, and on how much time they spent alone. The very happy students were more sociable and reported stronger friendships and romantic ties than the other groups, and good social relationships were present in essentially every member of the very happy group. None of them was happy all of the time.

Reading the research: A correlational, cross-sectional design: groups were selected on happiness scores, not assigned, and everyone was measured at a single point. Strong relationships look necessary for very high happiness in this sample, but the design cannot show they are sufficient or which came first. The limitations are one campus and self-report.

Apply it

A scatterplot puts weekly screen time on the horizontal axis and GPA on the vertical. The reported correlation is r = −0.32.

Direction: negative; more screen time goes with somewhat lower GPA. Strength: modest. The cloud slopes down but is wide, and plenty of heavy users have high GPAs. A plausible third variable is sleep, which could lower GPA and raise screen time at once. The defensible sentence: "In this sample, students reporting more weekly screen time tended to have slightly lower GPAs; these data cannot show that screen time causes lower grades." The failing version is "screen time hurts grades."

Exam tip: data-analysis items want direction and strength, then the conclusion the design rules out. Write both sentences.

Write it

Identify the research method used, and explain one conclusion these findings do not support.

Show a model response

This is a correlational study: the groups were defined by scores the students already had, and nothing was manipulated or randomly assigned. It does not support the conclusion that close relationships cause high happiness. The results fit equally well with happier people forming friendships more easily, or with a third variable such as extraversion raising both.

Why it earns the point: part A names the method, and the explanation identifies the directionality and third-variable problems instead of reciting a slogan.

Unit 2 review · 10 multiple-choice

Unit 2 review: Cognition

Ten questions in the three styles the exam uses (concept application, data analysis, and scientific investigation) one drawn from each lesson in the unit; click an option to see why that choice is right or wrong.

Multiple choice

  1. During a chemistry demonstration Priya is counting how many times the instructor stirs a beaker. She does not notice a classmate walk slowly across the front of the room holding a large yellow sign. Which of the following best explains what happened?

    Change blindness needs two views with something altered between them: a photograph that differs after a flicker, a person swapped during an interruption. Nothing here changed; an object was present the whole time and simply never registered.

    Sensory adaptation happens at the receptor level and requires constant, unchanging stimulation, like the smell of a room you have been sitting in for an hour. A classmate walking past is novel and moving, and Priya's retina responded to him normally.

    This is the finding behind the gorilla-in-the-video study: a demanding counting task consumes attention, and an unexpected object in plain view goes unreported. Her eyes worked; the bottleneck was attention, which is why "she wasn't paying attention" names the outcome rather than the mechanism.

    A perceptual set biases how an ambiguous stimulus is interpreted, as when you read a smudged letter as the one you expected. Priya did not misread the sign: she never saw it at all.

  2. Theo cancels a beach vacation after a week of news coverage of a shark attack, then drives eight hours to the mountains instead without a second thought, even though the drive carries the far higher risk of injury. Which of the following best explains his judgment?

    Shark attacks are rare and highway crashes are common, but coverage reverses how easily each comes to mind. The availability heuristic substitutes "how quickly can I recall an example" for "how often does this actually happen," and a week of coverage is what loaded his memory.

    Representativeness judges probability by resemblance to a stereotype: deciding someone is a librarian because he is quiet and tidy. Theo is not matching the beach to a category; he is over-weighting one event he can picture vividly.

    Confirmation bias would mean he already feared sharks and then went looking for stories that supported it. In the stem the coverage arrives first and the fear follows, which is availability rather than a biased search.

    Anchoring requires a specific number to anchor on: a listed price, a first offer, an initial estimate that later judgments drift toward. No figure appears anywhere in the stem.

  3. A pharmacist hears a 12-character prescription code read aloud once and, without writing it down, types it correctly ten seconds later. On the way to the keyboard she has recoded it as a four-digit year, a three-letter drug abbreviation she uses daily, and a five-digit lot number. Which of the following best explains her success?

    Iconic memory is a visual store that fades in well under a second, and she heard the code rather than seeing it. Even echoic memory, its auditory equivalent, lasts only a few seconds and could not carry her to the keyboard.

    Deep semantic processing builds durable long-term memories, which is not what happened here: ask her an hour later and the code will be gone. Recoding something for immediate use is not the same as encoding it for storage.

    Capacity limits are remarkably stable across people; experts do not get extra slots, they pack more into each slot. Assuming she has unusual capacity skips the mechanism the stem hands you.

    Chunking changes what counts as an item. Three familiar units sit well inside the roughly seven-item limit, and the phonological loop rehearses them for the few seconds she needs. This is why experienced people look like they remember more within their own field.

  4. Data: described bar graph. Participants answered one question about each of 60 words (about the word's appearance, its sound, or its meaning), and then took an unexpected recognition test.

    The vertical axis is the percentage of words later recognized, running from 0 to 80%, and there are three bars. The bar for appearance questions ("Is the word printed in capital letters?") reaches 18%. The bar for sound questions ("Does the word rhyme with train?") reaches 39%. The bar for meaning questions ("Does the word fit in this sentence?") reaches 63%. A note beneath the graph gives the mean time taken to answer each type of question: 0.6 seconds for appearance, 0.8 seconds for sound, and 1.1 seconds for meaning.

    Which statement is best supported by the data as presented?

    The answer times point the other way. Meaning questions took nearly twice as long as appearance questions, so processing time varies right alongside depth instead of being held constant, which is what a confound looks like.

    Both halves of this statement come straight off the display: 63 divided by 18 is about 3.5, and the note shows the deepest condition also consumed the most time. Naming the finding and then the alternative explanation it cannot exclude is exactly what a data-analysis item rewards.

    Repetition is held constant here: every word was presented once and only the question about it differed. A variable that does not vary cannot explain a difference in the results.

    Using the same words for every participant controls for word difficulty, which is a real strength, but it does nothing about the timing difference the note reports. Controlling one variable is not controlling all of them.

  5. A man who developed anterograde amnesia after damage to both hippocampi is taught to use an unfamiliar espresso machine. Across five daily sessions his time to produce a cup falls steadily from four minutes to under one. Asked each day whether he has used the machine before, he says he has never seen it. Which of the following best accounts for this pattern?

    Retrograde amnesia erases memories formed before an injury, and these sessions happen after it. His difficulty is with the new, which is precisely what anterograde amnesia means.

    Working memory holds information for seconds, not across a night's sleep. Nothing held in working memory could survive from one daily session to the next, so it cannot explain a five-day improvement.

    This dissociation is the classic finding from the patient known as H.M.: medial temporal damage blocks new explicit memories while procedural learning, supported by the cerebellum and basal ganglia, proceeds normally. He knows how without knowing that.

    Sensory adaptation is a decline in receptor response to constant stimulation; it does not make anyone faster at a sequence of actions. A steady gain across five sessions is learning, not a change in sensitivity.

  6. Described study

    Researchers tested 32 members of a university diving club. Each diver learned one list of 36 unrelated words while standing on a beach, and a second, equivalent list while submerged about fifteen feet underwater in scuba gear. Each diver was then tested for free recall twice: once in the setting where that list had been learned and once in the other setting. The order of lists and settings was counterbalanced across divers. Recall was substantially higher when the learning and testing settings matched than when they did not.

    Every diver completed all four learning-and-testing combinations. What does this feature of the design accomplish?

    This is a repeated-measures, or within-subjects, design. The comparison happens inside each person, so stable differences in memory ability, diving experience, and motivation are held constant instead of being left to chance.

    The match between learning and testing setting was manipulated by the researchers, and every diver met every combination, which makes this an experiment. Repeated measures is a way of assigning conditions, not a retreat from manipulating them.

    Generalization depends on who the participants are and how large the manipulation is, not on how conditions were assigned. A beach and open water differ far more than a dorm room and a classroom, so the everyday version of this effect is much smaller.

    Recall still has to be operationalized: here as the number of words from a 36-word list produced without cues. Measuring the same people repeatedly changes what the comparison controls, not whether the measure needs defining.

  7. Described study

    Researchers showed 90 adults at a driving school the same 30-second video of a collision between a car and a bicycle. Participants who attended the morning sessions were asked how fast the car was going when it smashed into the bicycle; those who attended the afternoon sessions were asked how fast it was going when it hit the bicycle. Everything else about the questionnaire was identical. Mean speed estimates were 41 mph in the morning group and 34 mph in the afternoon group. The researchers concluded that the verb changed the estimates.

    Which of the following is the most serious threat to that conclusion?

    A speed estimate is a serviceable operational definition: a number, produced under stated conditions, that another researcher could collect the same way. The weakness here is not the measure but what produced the difference between the groups.

    Random assignment is what makes two groups equivalent before a manipulation. Assigning by session time lets a whole bundle of differences (alertness, driving experience, who is free in the morning) ride along with the verb, so the 7 mph gap has more than one candidate explanation.

    This is a genuine limitation on generalizing to real witnesses, and it would affect the absolute level of both groups' estimates. But it applies equally to everyone, so it cannot account for the difference between the groups.

    A neutral-verb condition would make the finding richer by showing which direction the shift runs from. It would not repair a comparison in which the groups were formed by when people happened to show up.

  8. Naomi is a strong mathematics student. Just before a difficult placement test, the proctor remarks that the test is designed to measure whether students like her can handle advanced coursework. Naomi spends much of the period worrying about how her score will be read and finishes fewer items than usual. Which concept best explains her performance?

    Test bias is a property of a test's predictive accuracy across groups, established with data on how scores forecast later outcomes. Nothing in the stem concerns prediction; what changed was Naomi's state in the room.

    The Flynn effect is the long-run rise in raw test scores across generations, a fact about norms over decades. It has no bearing on one student's performance on one afternoon.

    A self-fulfilling prophecy runs through someone else's expectations changing their behavior toward the person: different teaching, different grading, fewer opportunities. The proctor scored nothing differently here; the effect ran through Naomi's own attention.

    Stereotype threat is the term the exam wants for this pattern: making a negative stereotype relevant adds a monitoring load, and working memory spent on worry is unavailable for the problems. The prediction is specific: the effect shows up on difficult items and disappears when the framing is removed.

  9. Data: summary statistics for the same 100-point final project in two sections of one course, each with 28 students.

    StatisticSection ASection B
    Mean7878
    Median7885
    Standard deviation414
    Lowest score7031
    Highest score8696

    Which statement is best supported by the table?

    When the mean sits below the median, the tail runs to the left: the extreme scores are low ones, and they drag the mean down while barely moving the middle value. A low score of 31 and a standard deviation of 14 are what produce that gap, and the median is the defensible summary.

    Positive skew means a long right tail, with the mean above the median. Section B shows the opposite ordering, mean 78 and median 85, and a wide range by itself tells you nothing about which direction the tail runs.

    Spread measures how alike the scores are, not how much anyone learned. Section B's median is seven points higher than Section A's, so the section with the tighter distribution is not the higher-scoring one.

    Equal means do not imply equally useful means. In Section A one standard deviation puts roughly two-thirds of students between 74 and 82; in Section B the same band runs from 64 to 92, which describes almost nobody.

  10. Data: correlations between four self-reported study habits and final exam score, from a survey of 180 students in one course. No variable was manipulated.

    VariableCorrelation with final exam score
    Hours per week spent self-testing+0.41
    Hours per week spent rereading notes+0.06
    Hours of sleep the night before the exam+0.22
    Hours per week at a part-time job−0.35

    Which statement is best supported by these data?

    An r of +0.06 is essentially no linear relationship, and its sign is positive in any case. Reading a near-zero coefficient as a negative relationship confuses a weak association with a reversed one.

    Strength is distance from zero, so +0.41 is the largest of the four, just ahead of −0.35. The second clause is the part the exam is really testing: in a survey with no manipulation, students who self-test more may also be the students who attend more, sleep more, or care more.

    Direction and strength are separate readings. The sign says which way the relationship runs; the distance from zero says how strong it is, and 0.35 is larger than 0.22, so the job variable is the stronger of the two.

    This is a causal claim drawn from correlational data, with a prescription stacked on top of it. Work hours may stand in for financial pressure, commuting time, or lost sleep, and cutting the hours would not remove those.

Lesson 3.1 · Unit 3 · CED topic 3.1

Themes and methods in developmental psychology

Developmental psychology asks how people change across a whole life. That creates a problem no other subfield has: the variable you care about is time, and nobody can assign it.

Three arguments run under the unit: smooth growth or stages, stability or change, body or experience.

Key terms
  • Cross-sectional design: several age groups compared at one moment; fast, but they differ in more than age.
  • Longitudinal design: the same people measured repeatedly for years; it shows change within a person, and loses people to attrition.
  • Cohort effect: a difference produced by the era a group grew up in, not by age.
  • Sequential design: several cohorts each followed over time, so age and cohort pull apart.
  • Continuity vs. discontinuity; stability vs. change: smooth growth or stages; a trait that persists or shifts.
  • Maturation, critical and sensitive periods, teratogen: biologically timed growth; windows when experience matters most; an agent harming prenatal development.
Study spotlight

Study: Moffitt and colleagues (2011)

In 1972 and 1973 researchers enrolled 1,037 babies born in one New Zealand city and have assessed them repeatedly since, from age 3 into their forties, with retention above 90% at most waves. Parents, teachers, and observers rated each child's self-control through childhood; decades later those ratings were matched against adult records of health, finances, and criminal convictions. Children lowest in self-control became adults with poorer health, more money trouble, and more convictions, as a gradient across the whole range.

Reading the research: A longitudinal correlational study. Self-control was operationalized as childhood ratings, the outcomes as adult medical, financial, and court records, so the time order is clean, but nothing was manipulated, and families differed in income too. The limits are cost, attrition, and one cultural cohort.

Apply it

Design two studies of vocabulary growth from ages 4 to 10.

Cross-sectional: test 4-, 6-, 8-, and 10-year-olds and compare means. Fast, but the oldest group learned to read under a different curriculum, so a cohort effect masquerades as growth. Longitudinal: follow one group of four-year-olds until they turn ten. The change is now within the same children, at the cost of six years and the families who leave.

Exam tip: a scientific investigation item hands you a design and asks what it cannot conclude.

Write it

Identify the research method used in this study, and explain one limitation of claiming that low self-control causes adult health problems.

Show a model response

The researchers used a longitudinal correlational study: one group was measured over decades and nothing was manipulated. Because children were not assigned to high or low self-control, a third variable such as family income could produce both the ratings and the outcomes, so the link is not causal.

Why it earns the point: it names the method instead of retelling the procedure, then supplies a specific confound.

Lesson 3.2 · Unit 3 · CED topic 3.2

Physical development across the lifespan

The sequence of physical development is close to fixed; the timing is not. Knowing which parts run on a biological clock and which wait on experience is what separates ordinary variation from a delay worth checking.

Key terms
  • Zygote, embryo, fetus: the fertilized cell for two weeks, weeks three to eight while organs form, then week nine to birth.
  • Teratogen: an agent that crosses the placenta and harms development; prenatal alcohol exposure can produce fetal alcohol spectrum disorder.
  • Newborn reflexes: rooting, sucking, grasping, and the Moro startle; they fade as the cortex matures.
  • Motor milestones: rolling, sitting, crawling, walking, in a fixed order on a variable clock.
  • Synaptic pruning and myelination: unused connections are cut and axons gain insulation, so brains speed up by losing.
  • Puberty and after: primary and secondary sex characteristics appear; the prefrontal cortex matures into the twenties, after the reward systems it restrains; menopause ends the menstrual cycle at midlife.
Study spotlight

Study: Gibson & Walk (1960)

Gibson and Walk built a table with a glass top. Under one half a checkered surface lay directly against the glass; under the other half the same pattern lay several feet below, so the glass looked like a drop. Thirty-six infants aged 6 to 14 months were placed on the center board while a caregiver called from one side. Twenty-seven crawled across the shallow side. Only three ventured onto the glass over the apparent cliff; the rest backed away or cried.

Reading the research: A controlled observation whose two sides were identical except for apparent depth; the dependent variable was whether the infant crossed. The limitation is who could be tested: only infants already crawling, so it cannot say whether depth perception is present at birth.

Apply it

A parent asks why their fearless nine-month-old suddenly refuses the top step.

Two things changed together. Maturation supplied the visual system and muscle control to crawl, and crawling supplied experience with edges: the infants most likely to refuse the deep side are the ones who have been moving on their own longest. The caution is new competence, not new fragility.

Exam tip: a concept application item rewards naming the interaction of maturation and experience, not just "he grew up."

Write it

Explain the extent to which these findings generalize, citing specific evidence about the participants.

Show a model response

They generalize only so far. The participants were 36 human infants between 6 and 14 months old, all of whom could already crawl, because crossing was the measure. The findings describe infants with months of self-produced movement; they cannot tell us what a newborn perceives about depth.

Why it earns the point: part E wants participant evidence, so it names the age range and the crawling requirement.

Lesson 3.3 · Unit 3 · CED topic 3.3

Gender and sexual orientation

Everyday talk bundles three things together that psychology pulls apart: what a body is, what a culture expects, and whom a person is drawn to. They vary independently, and precise, respectful terms are part of the science rather than a courtesy added afterward.

Key terms
  • Sex and gender: sex refers to biological characteristics: chromosomes, hormones, anatomy; gender to the roles a culture attaches to them.
  • Gender identity: a person's own sense of their gender, which may or may not match the sex assigned at birth.
  • Gender role and gender typing: the behavior a culture expects, and how a child takes it on.
  • Androgyny: endorsing traditionally masculine and traditionally feminine traits about equally, rather than one set.
  • Gender schema theory and social learning: children build a framework for gender and sort the world into it, and they imitate models who are reinforced.
  • Sexual orientation and the sexual response cycle: an enduring pattern of attraction; the excitement, plateau, orgasm, and resolution phases.
Study spotlight

Study: Bem (1974)

Bem set out to measure masculinity and femininity as two separate dimensions rather than opposite ends of one scale. Undergraduates rated how desirable each of hundreds of personality traits was in American society for a man, or for a woman, not how well it described them. Twenty traits rated far more desirable for men, twenty for women, and twenty rated neutral became the inventory. When about a thousand students then rated themselves on those items, roughly a third scored androgynous.

Reading the research: Scale construction and survey research, not an experiment: nothing was manipulated, so no causal claim follows. The scoring key rests on the raters' judgments, which is the limitation: the norms record what one culture in one decade counted as masculine or feminine.

Apply it

A student wears clothing her school codes as masculine, says that she is a woman, and has dated only women. Sort the three details.

The clothing speaks to gender role: behavior a culture codes, not an inner state. "She is a woman" is gender identity, reported by the person herself, and it fixes the words to use for her. Dating only women is sexual orientation. The three vary independently, so none predicts the others.

Exam tip: distractors trade these three. Read for whose report it is.

Write it

Identify the research method used in this study, and explain one limitation on generalizing its norms.

Show a model response

Bem used survey research to build a scale: students rated traits and then rated themselves, and no variable was manipulated. The limitation is who supplied the norms: American college students in the early 1970s, whose judgments became the scoring key, so the scale travels poorly to other cultures or later decades.

Why it earns the point: it names the method rather than the topic, then ties the limitation to a specific feature of the sample.

Lesson 3.4 · Unit 3 · CED topic 3.4

Cognitive development across the lifespan

Children are not small adults with less information. For a while they reason by rules that are internally consistent and simply wrong, and the errors are the evidence: they show what a mind can and cannot yet do.

Key terms
  • Schema, assimilation, accommodation: a mental framework; fitting new input into one; changing it when input will not fit.
  • Sensorimotor stage (birth to about 2): knowing through action, ending with object permanence: hidden things still exist.
  • Preoperational stage (about 2 to 7): symbols and pretend play, alongside egocentrism, failure to conserve, and a theory of mind still forming.
  • Concrete then formal operational: logic about concrete things from about 7; hypothetical and abstract reasoning from about 12.
  • Zone of proximal development, scaffolding, private speech: what a learner can do with help, support that fades, self-talk that becomes thought.
  • Dementia: later-life decline in memory and reasoning severe enough to disrupt daily life; not ordinary aging.
Study spotlight

Study: Piaget (1950s); Baillargeon (1987)

Piaget worked with children one at a time, following each answer with questions about why. In the conservation task a child watched water poured from a short wide glass into a tall narrow one, then was asked whether there was more, less, or the same. Children under about seven typically said the tall glass held more; older children said the amount was unchanged, since nothing was added or removed. Baillargeon later tested infants by timing how long they looked at physically impossible events; infants as young as 3½ months looked longer at the impossible outcome.

Reading the research: Piaget used the clinical interview: small samples, flexible follow-up questions, description rather than controlled comparison. Baillargeon's violation-of-expectation studies used looking time as the dependent variable and put object permanence months earlier. The limitation is the measure: a task needing reaching or talking hides knowledge a child already has.

Apply it

Predict what a four-year-old and a nine-year-old say about the poured water.

The four-year-old says the tall glass has more: he centers on height and cannot mentally reverse the pour; failure to conserve, the signature of preoperational thought. The nine-year-old says it is the same because nothing was added and you could pour it back. That reversibility is concrete operational reasoning.

Exam tip: name the stage and the operation the child lacks; "she is too young" earns nothing.

Write it

A preschool teacher says a four-year-old who hides by covering his own eyes is "being silly." Explain the behavior with a concept from this lesson.

Show a model response

This is egocentrism. A preoperational child has trouble taking another person's visual perspective, so he assumes that because he cannot see you, you cannot see him. It is a limit on perspective-taking, not silliness or deception, and it fades as theory of mind develops.

Why it earns the point: it names the concept, ties it to the specific behavior, and says what follows.

Lesson 3.5 · Unit 3 · CED topic 3.5

Communication and language development

No one teaches a toddler grammar, yet the order in which it arrives is nearly identical across languages. The puzzle is how much of that schedule a child brings and how much the people talking to her supply.

Key terms
  • Phoneme and morpheme: the smallest sound unit that changes meaning; the smallest unit that carries meaning, such as -s or -ed.
  • Semantics, syntax, grammar: what words mean, the rules for ordering them, and the whole system together.
  • Cooing, babbling, one-word stage, telegraphic speech: vowel sounds, then consonant-vowel strings, then single words near a year, then two-word utterances without function words.
  • Overgeneralization: applying a regular rule where the language is irregular ("goed," "foots").
  • Critical period and the language acquisition device: a window when a first language is learned readily; Chomsky's inborn readiness for grammar, against imitation and reinforcement.
  • Broca's and Wernicke's aphasia: effortful speech with meaning intact; fluent speech that carries little meaning, with impaired comprehension.
Study spotlight

Study: Berko (1958)

Berko showed children cartoon drawings of invented creatures and gave them spoken prompts to complete. Shown one bird-like figure, a child was told it was a wug; shown two of them, the child had to finish a sentence that called for the plural. Because the words did not exist, no child could have heard that plural before. Preschoolers of about four and five and first-graders were tested. About three-quarters of the preschoolers supplied "wugs," and nearly all first-graders did, so they were applying a rule rather than retrieving a memorized word.

Reading the research: A structured experimental task whose key control is the nonsense word: prior exposure cannot explain a correct answer, and the dependent variable is the form the child produces. The limitation is scope: English-speaking children in one region, and only a few morphological rules.

Apply it

A three-year-old who used to say "went" now says "I goed to the store."

This is progress. Earlier he produced "went" as a whole chunk copied from adults. Now he has extracted the rule, add -ed for the past tense, and applies it everywhere, including to a verb English makes irregular. Overgeneralization is the visible evidence of a rule being learned; the exceptions come later.

Exam tip: an application item hands you a child's error and asks what it shows. The answer is the rule underneath it.

Write it

Identify the research method used in this study, and explain how the use of invented words supports the conclusion.

Show a model response

Berko used an experimental task in which every child received the same controlled prompts. The invented words matter because a child cannot have memorized "wugs," so producing the plural must come from a rule applied to new material. That rules out imitation as the explanation.

Why it earns the point: it names the method, then explains what the control rules out instead of restating the finding.

Lesson 3.6 · Unit 3 · CED topic 3.6

Social-emotional development across the lifespan

A baby arrives with a reactive style already in place and starts building relationships on top of it. The question is what those relationships look like when they work, and how anyone measures something as vague-sounding as a bond.

Key terms
  • Temperament: a biologically based style of reacting, visible in infancy and fairly stable.
  • Attachment patterns: secure; insecure-avoidant; insecure-anxious (resistant); and disorganized, which shows no consistent strategy.
  • Separation anxiety and contact comfort: distress when a caregiver leaves; Harlow's infant monkeys clung to a soft cloth surrogate over the wire one that fed them.
  • Parenting styles: authoritative (high demands, high warmth), authoritarian (high, low), permissive (low, high), neglectful (low, low).
  • Erikson's psychosocial stages: eight conflicts across a life; adolescence is identity vs. role confusion.
  • Adolescent egocentrism and ecological systems: the imaginary audience and personal fable; nested layers of influence from family out to culture.
Study spotlight

Study: Ainsworth and colleagues (1970, 1978)

About a hundred one-year-olds were observed in an unfamiliar playroom through eight episodes of roughly three minutes: the caregiver present, a stranger entering, the caregiver leaving and returning, twice. Trained coders scored what the infant did at reunion, not during the separation. Roughly two-thirds explored while the caregiver was present, protested when she left, and were comforted on her return: the pattern labeled secure. The rest divided between avoidant infants, who ignored the returning caregiver, and resistant infants, who sought contact and then resisted it.

Reading the research: Structured observation with trained coders. The dependent variable is coded reunion behavior; nothing is manipulated, so it describes rather than explains. The limitation: a brief laboratory episode, and categories built in one country, travel imperfectly to cultures with different norms about separation.

Apply it

A toddler cries when her mother leaves, crawls straight to her on her return, settles in under a minute, and goes back to the blocks.

That is the secure pattern: distress at separation, comfort accepted at reunion, a return to exploring, the caregiver used as a secure base. Then say what the label does not do. It describes one relationship on one day, predicts group averages rather than one child's future, and can differ with another caregiver.

Exam tip: classification items reward the reunion detail, not the crying.

Write it

About two-thirds of the infants were classified as secure. Describe what that proportion does and does not indicate about attachment.

Show a model response

It indicates that secure was the most common classification in this sample: most one-year-olds observed used the caregiver as a base for exploring and were comforted at reunion. It does not indicate that the remaining third had poor relationships, and as a proportion from one American sample it does not estimate how attachment is distributed everywhere.

Why it earns the point: part C asks what a statistic indicates, so it reads the proportion and then bounds the claim.

Lesson 3.7 · Unit 3 · CED topic 3.7

Classical conditioning

Some learning takes no effort and no intention. A signal that reliably precedes something important starts producing the reaction on its own, which is why a ringtone can tighten your chest before you read the name on the screen.

Key terms
  • UCS and UCR: a stimulus that triggers a reaction without any learning, and that automatic reaction.
  • NS, CS, and CR: a neutral stimulus becomes conditioned by pairing with the UCS, then triggers the learned response.
  • Acquisition, extinction, spontaneous recovery: pairings build the response; the CS alone fades it; a rest brings part of it back.
  • Generalization and discrimination: responding to stimuli resembling the CS; learning to respond only to the CS.
  • Higher-order conditioning: an established CS can itself condition a second neutral stimulus.
  • Taste aversion, preparedness, counterconditioning: one flavor-illness pairing can produce lasting avoidance, since some associations come more readily; pairing the CS with an incompatible response undoes it.
Study spotlight

Study: Watson & Rayner (1920)

An eleven-month-old infant was shown a white rat and reached for it without distress. The researchers then struck a steel bar behind his head as he touched the rat; the clang startled him and he cried. After a handful of pairings across two sessions the rat alone was enough: he cried and crawled away. He also showed distress toward a rabbit and a fur coat, reported as generalization. Nothing was done afterward to undo the conditioning.

Reading the research: A single-participant demonstration, not a controlled experiment: one infant, no control condition, fear measured by the researchers' own observation. It could not run today: an IRB would refuse to induce fear in an infant with no benefit to him and no plan to extinguish it.

Apply it

A patient's heart races at the whine of a dental drill. Label the parts, then plan a fix.

The UCS is the pain of the drill reaching a nerve; the UCR is the pain reaction. The sound was neutral until it preceded that pain; it is now the CS, and the racing heart is the CR. To countercondition, pair recordings of the sound with deep relaxation, quiet enough at first to produce no anxiety and louder only while she stays calm.

Exam tip: find the UCS first, the CS is whatever used to be neutral.

Write it

A student stung by a wasp beside a humming vending machine now feels fear at the hum. Identify the CS and CR, and predict what a month of daily, sting-free use does.

Show a model response

The CS is the hum and the CR is the fear it triggers; the sting was the UCS. Presenting the CS without the UCS is extinction, so the fear should weaken over the month. It may not vanish, after a break, spontaneous recovery can bring a weaker version back.

Why it earns the point: it labels both terms correctly and names extinction and spontaneous recovery instead of saying the fear "goes away."

Lesson 3.8 · Unit 3 · CED topic 3.8

Operant conditioning

Classical conditioning explains reactions that happen to you. Operant conditioning explains what you do on purpose: behavior is shaped by what follows it, and the timing of consequences matters more than their size.

Key terms
  • Law of effect: behavior followed by a satisfying consequence is repeated; the principle Skinner built on.
  • Reinforcement and punishment: reinforcement raises a behavior's rate, punishment lowers it; positive adds a stimulus, negative removes one.
  • Primary and secondary reinforcers: food and relief work without learning; money, grades, and tokens acquire value by association.
  • Shaping: reinforcing successive approximations until the whole behavior appears.
  • Schedules: fixed- or variable-ratio after a set or varying number of responses; fixed- or variable-interval after a set or varying time. Variable-ratio gives the steadiest rate and the most resistance to extinction.
  • Extinction burst, token economy, instinctive drift: a spike when reinforcement stops; tokens exchanged for rewards; species-typical behavior intruding on a trained sequence.
Study spotlight

Study: Skinner (1948)

Eight hungry pigeons were placed in cages where a food hopper opened for a few seconds every fifteen seconds, no matter what the bird was doing. Within days, six of the eight had developed distinctive repeated actions between deliveries: one turned counterclockwise, one thrust its head into a corner, another swung its head like a pendulum. Skinner's reading was that whatever a bird happened to be doing when food appeared was strengthened, so the more it repeated that action, the more often food happened to coincide with it.

Reading the research: An animal experiment observed within subjects; the independent variable is the delivery schedule and the dependent variable is the behavior coded between deliveries. Limitations: eight birds, "ritual" judged by an observer, and a later reinterpretation in terms of species-typical food-anticipation behavior.

Apply it

A game sells loot boxes whose contents are random. Name the schedule and say why quitting is hard.

It is a variable-ratio schedule: reward comes after an unpredictable number of purchases. That gives a high, steady rate and the most resistance to extinction, because no run of empty boxes is unusual enough to signal the payoff has stopped. A paycheck every two weeks is fixed-interval: effort rises near payday, and the pattern is easy to break.

Exam tip: ratio counts responses, interval counts time; variable means unpredictable.

Write it

A teacher drops a daily homework check after a student turns work in on time three days running, and his on-time rate climbs. Name the procedure and justify the label.

Show a model response

This is negative reinforcement. Something the student found unpleasant, the daily check, was removed, and the behavior it followed, turning work in on time, became more frequent. It is reinforcement rather than punishment because the rate went up, and negative because a stimulus was taken away.

Why it earns the point: it settles the label with the two questions the exam rewards; did the behavior increase, and was something added or removed?

Lesson 3.9 · Unit 3 · CED topic 3.9

Social, cognitive, and neurological factors in learning

Conditioning cannot explain learning you never performed and were never rewarded for. Watching is enough, and so is thinking: something happens between stimulus and response that a strictly behaviorist account leaves out.

Key terms
  • Observational learning, modeling, vicarious reinforcement: acquiring behavior by watching a model, with the model's rewards or punishments shifting how likely imitation is.
  • Mirror neurons: cells that fire when an animal acts and when it watches the same act; how much they explain about imitation is debated.
  • Latent learning and cognitive maps: learning that stays hidden until there is a reason to use it; Tolman's rats had mapped the maze before any reward.
  • Insight learning: a solution that arrives whole after a pause, not by gradual trial and error.
  • Intrinsic and extrinsic motivation; overjustification: doing something for its own sake or for a reward; paying for an already-enjoyed activity can undercut interest.
  • Learned helplessness: after aversive events nothing could control, an organism stops trying even once escape becomes possible.
Study spotlight

Study: Bandura, Ross & Ross (1961)

Seventy-two children at a university nursery school, half boys and half girls, were brought one at a time into a room with an adult. Some watched the adult play quietly; others watched the adult strike an inflatable doll in specific ways, with a mallet and with set phrases. Each child was then mildly frustrated and left alone with toys including the doll, while observers counted behaviors through a one-way mirror. Children who had seen the aggressive model produced far more imitative aggression, and boys imitated physical aggression more than girls did.

Reading the research: A laboratory experiment. The independent variable is the model's behavior; the dependent variable is the count of coded aggressive acts. The limitation is what was measured: hitting an inflatable toy built to be hit is not harming a person, and the delay was minutes.

Apply it

Evaluate the claim "violent video games cause violence," using this study.

Observational learning and vicarious reinforcement do predict imitation, especially when the model is rewarded. But say what the design licenses: a short-term rise in imitative aggression toward a toy, minutes later, in a laboratory. Real violence needs evidence this design cannot give: long delays, human targets, and control for who chooses the games.

Exam tip: an argument item wants the concept and the boundary of the evidence.

Write it

A coach starts paying players $5 per practice. Attendance rises for a month, then drops below where it began once the payments stop. Name the effect and explain it.

Show a model response

This is the overjustification effect. The players already practiced for their own reasons, and the $5 supplied an extrinsic reason that came to explain the behavior instead. When the reward stops, that reason is gone and the original interest has weakened, so attendance falls below where it began.

Why it earns the point: it names the effect and traces the mechanism from intrinsic to extrinsic motivation rather than just saying rewards backfire.

Lesson 3.10 · Unit 3 · Research methods

Effect size, significance, and replication

A finding is not a fact until it survives being run again. Two questions decide how much weight a result carries: how big the difference is, and whether anyone else can produce it.

Key terms
  • Effect size (Cohen's d): the difference between two means in standard deviations; about 0.2 is small, 0.5 medium, 0.8 large.
  • Statistical vs. practical significance: a small p-value says a result would be unlikely if nothing were going on; it does not say the effect is big or useful.
  • Direct and conceptual replication: running the same procedure again; testing the same idea a different way.
  • Publication bias and the file drawer: journals favor positive results, so null findings go unpublished and the literature looks stronger than the evidence.
  • p-hacking: trying analyses until one crosses the threshold, which manufactures significance out of noise.
  • Preregistration and meta-analysis: posting the hypothesis and analysis plan before collecting data; pooling many studies into one effect-size estimate.
Study spotlight

Study: Baumeister and colleagues (1998); Hagger and colleagues (2016)

A few dozen undergraduates asked to skip a meal were seated at a table holding fresh-baked cookies and a bowl of radishes, and told which they were allowed to eat. Everyone then worked on puzzles that had no solution, and the measure was how long they kept trying. Those made to resist the cookies gave up after about 8 minutes; those who ate the cookies persisted about 19. Nearly twenty years later, 23 laboratories ran a preregistered replication with over 2,000 participants on one agreed procedure, and found an effect indistinguishable from zero, about d = 0.04.

Reading the research: Both are experiments with random assignment. The original pairs a striking difference with a small sample; the replication pairs a tiny difference with an enormous one. A single significant result is an invitation to replicate, not a conclusion.

Apply it

Two studies measure the same outcome. One reports d = 0.2, the other d = 0.8. Which matters more?

The second. A d of 0.8 puts the average treated person about eight-tenths of a standard deviation above the average control; a d of 0.2 barely separates the distributions. A small p-value does not settle it, because p depends on sample size as well as effect: with enough participants, the d = 0.2 study can report the smaller p-value.

Exam tip: when an item gives you a p-value, ask what the effect size was.

Write it

Describe what an effect size of about d = 0.04 in the replication indicates about ego depletion.

Show a model response

It indicates that the depleted and control groups differed by about four-hundredths of a standard deviation: close to no difference at all. Because the replication pooled more than 2,000 participants across 23 laboratories, the estimate is precise, so an effect as large as the original is unlikely, though a very small one cannot be ruled out.

Why it earns the point: part C asks what a statistic indicates, so it translates d into standard deviations and says what the precision allows.

Unit 3 review · 10 multiple-choice

Unit 3 review: Development and Learning

Ten questions in the three styles the exam uses (concept application, data analysis, and scientific investigation) one drawn from each lesson in the unit; click an option to see why that choice is right or wrong.

Multiple choice

  1. Described study

    To study how memory for names changes with age, researchers recruit 90 volunteers in three groups: 30 people aged 20 to 29, 30 aged 50 to 59, and 30 aged 75 to 84. All are tested on the same afternoon in the same room. Each learns twenty face-and-name pairs and is tested ten minutes later. The mean number of names recalled is 14.2 in the youngest group, 11.6 in the middle group, and 8.3 in the oldest group. The researchers conclude that memory for names declines with age.

    Which feature of this design most limits that conclusion?

    A ten-minute delay limits what the study can say about long-term retention, but every group faced the same delay. A feature shared by all three conditions cannot explain the differences among them.

    Cross-sectional designs compare different people at a single moment, so everything that separates someone born in 1945 from someone born in 2000 travels with the age difference. A longitudinal or sequential design is what pulls age and cohort apart.

    Thirty per group is a workable sample for a comparison of this kind, and the reported differences are large. Sample size affects how precise an estimate is; it is not the structural problem in this design.

    Participants cannot be randomly assigned to an age: age is a participant variable, not something a researcher can hand out. That is exactly why developmental research so often has to settle for quasi-experimental comparisons.

  2. Jonas, 16, can explain in detail why texting while driving is dangerous and scores well on the written portion of his driving test. With three friends in the car on a Friday night, he texts anyway. Which of the following best explains the gap between what he knows and what he does?

    The developmental mismatch is the point. Reasoning about risk is fully available to him, the written test proves it, but the systems that make a reward feel urgent, particularly a social reward, come online years before the frontal machinery that brakes them.

    Formal operational reasoning is precisely what he demonstrates when he explains the danger and passes the written test. His difficulty is not thinking about hypotheticals; it is which system wins in the moment.

    Synaptic pruning begins in early childhood and continues prominently through adolescence, so it is not a process still waiting to start. The raw number of synapses is not what settles a single decision either.

    The personal fable, the sense of being uniquely protected, is a real feature of adolescent egocentrism and may well contribute here. But "literally unable to perceive" overstates it: he perceives the risk and discounts it, which is the distinction this item is testing.

  3. Described study

    A gender-role inventory built in the early 1970s was scored from ratings by American undergraduates of how desirable each of hundreds of traits was for a man, or for a woman, in their society. Twenty traits rated far more desirable for men, twenty far more desirable for women, and twenty rated neutral became the scale. A researcher now administers that inventory, with its original scoring key, to 800 undergraduates at the same university and finds that far fewer students score in the masculine range than the original norms recorded.

    Which interpretation of the new result is most defensible?

    This treats the scale as a fixed yardstick, when the yardstick itself was built out of one era's ratings. A trait that counted as masculine in 1974 may simply no longer be sorted that way, which would move scores without anyone changing at all.

    Test-retest reliability concerns whether the same people score similarly across a short interval, not whether two different samples fifty years apart score alike. Genuine change over decades is not evidence that a measure is inconsistent.

    Neither sample was drawn randomly from all adults, yet comparing two student samples from the same university is still informative. The caution belongs on what the scale's norms mean, not on whether any comparison can be made.

    This is the limitation the exam expects you to name for scale-construction and survey research: the norms are cultural data. The result is genuinely interesting and genuinely ambiguous between a change in people and a change in what the culture codes as masculine, and this design cannot separate the two.

  4. A five-year-old watches as ten pennies in a tight row are spread out into a longer row. He counts ten pennies before the change and ten after it, then insists the longer row now has more. Which of the following best describes his reasoning?

    Object permanence, knowing a hidden object still exists, is typically in place before a child's second birthday, and these pennies were never hidden. He can see all ten of them the entire time.

    Accommodation means changing a schema when new information will not fit it, which is the thing he is failing to do. He is assimilating the spread-out row to a rule he already holds: longer looks like more.

    Conservation is the understanding that quantity is unchanged by a change in appearance. Centering on a single dimension and being unable to mentally undo the transformation are the two limits Piaget identified, and both show in his answer even though his counting is accurate.

    Egocentrism is difficulty taking another person's perspective, as when a child hides by covering his own eyes. Nothing in this task turns on what he believes the adult can see.

  5. At two, Lena said "I ran" and "two mice." At three and a half she says "I runned" and "two mouses." Her grandfather worries that her language is going backwards. Which of the following is the best response?

    Telegraphic speech is the two-word stage, "want cookie," "Daddy go", in which endings and function words are missing altogether. Lena is adding endings, not dropping them.

    Earlier, "ran" and "mice" were whole items copied from adults. Now she has abstracted the rules (add -ed, add -s), and applies them everywhere, including where English is irregular. The error is the evidence for the rule, which is why it appears after the correct forms rather than before them.

    Critical and sensitive periods describe windows during which experience has an outsized effect, such as the early years for a first language. The label does not explain why one particular error emerges, and "they disappear on their own" is not a mechanism.

    She produces both the -ed and the -s ending accurately, so articulation is not the problem, and she says "runned" rather than mispronouncing "ran." The error is grammatical, not phonological.

  6. Data: 100 one-year-olds observed in a laboratory playroom through a standard sequence of separations from and reunions with a caregiver. Trained coders classified each infant's reunion behavior and recorded two timed measures.

    ClassificationSecureAvoidantResistant
    Number of infants652114
    Mean seconds exploring while caregiver present14211863
    Mean seconds of contact-seeking at reunion389165

    A reader concludes from the table that the avoidant infants were the least distressed by the separation. Which response is best supported by the data?

    Contact-seeking is a behavior, not a feeling, and the whole point of an operational definition is that the two are not interchangeable. An infant can be highly aroused and still not approach the caregiver.

    Exploration while the caregiver is present speaks to whether the infant uses her as a secure base, and it is measured before the separation even happens. It cannot establish what the separation did.

    The table shows resistant infants exploring the least, at 63 seconds, with avoidant infants at 118. That column does not contradict the reader's claim: it simply does not address it.

    This is a "what the data do not show" item, and the answer sits in the row labels. Neither timed measure records separation distress, so the reader has quietly swapped a coded reunion behavior for an internal state the study never measured.

  7. A patient who is receiving chemotherapy eats a lemon drop in the waiting room before each infusion. The drug reliably causes nausea a few hours later. After four rounds, the smell of lemon alone makes her feel sick. In this example, what is the conditioned stimulus?

    Nausea produced directly by the drug is the unconditioned response: an automatic reaction that required no learning. A response is never a stimulus in this scheme, so it cannot be the conditioned stimulus.

    The lemon was neutral before the pairings and now triggers the reaction on its own, which is the definition of a conditioned stimulus. Taste aversions like this one form remarkably fast, often after a single pairing, because organisms are biologically prepared to link flavours with illness.

    The drug is the unconditioned stimulus: it produces nausea with no learning history at all. Confusing the unconditioned stimulus with the conditioned stimulus is the most common error on these items, and the test is always which one was neutral to begin with.

    Her queasiness at the smell of lemon is the conditioned response: the learned reaction to the conditioned stimulus. It looks much like the unconditioned response, but it is triggered by the signal rather than by the drug.

  8. Priya checks her email roughly every four minutes throughout the day. Messages she cares about arrive a few times a day at unpredictable times, and checking more often does not make one arrive any sooner. Which reinforcement schedule best describes her checking?

    A fixed-ratio schedule pays off after a set number of responses (every tenth sale, every fifth punch on a coffee card), and produces a burst-and-pause pattern. Nothing in the stem counts her checks.

    Variable-ratio is the slot-machine schedule, where responding faster genuinely earns more payoffs. The stem rules it out directly: checking more often does not make a message arrive sooner, so the reward is not tied to her response count.

    The question that separates these schedules is whether time or responses control availability. A message becomes available as time passes, and her checks matter only in that one of them eventually finds a waiting message: the signature of a variable-interval schedule, which produces a steady, moderate rate of responding.

    Fixed-interval means reinforcement becomes available after a set amount of time, which produces the scalloped pattern of a student who studies harder as a weekly quiz approaches. Her messages arrive unpredictably, so the interval is variable rather than fixed.

  9. Data: described bar graph. Preschoolers who already chose to draw during free play were randomly assigned to three conditions. In the expected-reward condition each child was promised a certificate for drawing; in the unexpected-reward condition each child received the same certificate afterward with no prior promise; in the no-reward condition no certificate was mentioned or given.

    The vertical axis is the percentage of free-play time spent drawing, running from 0 to 25%, and the horizontal axis has three pairs of bars, one pair per condition. The lighter bar in each pair is the baseline measured before the conditions began and the darker bar is the measurement taken two weeks afterward, with no certificates available. Expected reward: baseline 17%, follow-up 9%. Unexpected reward: baseline 17%, follow-up 17%. No reward: baseline 16%, follow-up 18%.

    Which conclusion is best supported by the graph?

    The comparison that matters is expectation, not reward: both rewarded groups received the identical certificate, and only the group promised it beforehand dropped. That isolates the mechanism: a promised reward supplies an external reason for an activity the child was already choosing freely.

    The unexpected-reward group received a certificate and stayed at 17%. Since one rewarded group declined and the other did not, "rewards reduce motivation" is contradicted by the graph's own data.

    A two-point move from 16% to 18% in a group where nothing was done is the size of change that ordinary variation produces. Reading it as an effect of withholding rewards treats noise as a finding.

    If the certificate were simply too small a reward, the unexpected-reward group, which received exactly the same certificate, should have declined as well. It did not, so the value of the reward cannot be what separates the groups.

  10. Data: four studies of the same tutoring program, each measuring the same 100-point reading test. Within-group standard deviations were about 10 points in all four studies.

    Study1234
    Participants (n)401,200602,400
    Mean difference (points)8.21.84.50.5
    Cohen's d0.820.180.450.05
    p.01.002.09.22

    Which statement is best supported by the table?

    A p-value answers a different question from an effect size: how surprising this result would be if nothing were going on, given this sample size. With 1,200 participants even a d of 0.18 clears any conventional threshold, while the effect itself is the second smallest in the table.

    A p-value above .05 means the result did not reach the conventional threshold, not that the effect is zero. Study 3's d of 0.45 is a moderate difference measured with only 60 participants, so the honest reading is "underpowered," not "no effect."

    Study 1 pairs the largest d with the smallest sample, which is exactly the combination that should slow you down. With 40 participants the estimate is imprecise, and publication bias means the literature over-represents large effects that came from small samples.

    Dividing the mean difference by the standard deviation is what makes the four studies comparable: 1.8 points on a scale with a standard deviation of 10 is d = 0.18, a gap most students would never notice. This is the practical-versus-statistical significance distinction the exam tests, and the fix is to read d before you read p.

Lesson 4.1 · Unit 4 · CED topic 4.1

Attribution theory and person perception

Watch someone cut a line and you do not think "long day, probably late." You think "rude." Social psychology's founding observation is that people explain behavior by reaching for the actor's character and skipping the pressures acting on them, and that the skip is systematic, not random.

Key terms
  • Attribution: dispositional vs. situational: an explanation that credits the person's traits, versus one that credits the circumstances they were in.
  • Fundamental attribution error and actor-observer bias: observers over-weight disposition when explaining others; for their own behavior the same people favour the situation.
  • Self-serving bias, just-world hypothesis, explanatory style: taking credit for successes and blaming circumstances for failures; assuming people get what they deserve; a stable optimistic or pessimistic habit of explaining events.
  • False consensus and halo effects: overestimating how many people share your view; letting one good quality color every other judgment of a person.
  • Social comparison and relative deprivation: judging yourself against others, and the discontent of measuring yourself against someone better off.
  • Stereotype, prejudice, discrimination: a belief about a group, an attitude toward it, and the behavior. Add implicit attitudes (automatic, often unendorsed), in-group bias, out-group homogeneity, and ethnocentrism.
Study spotlight

Study: Ross, Amabile & Steinmetz (1977)

Students were paired for a quiz game and assigned roles by a coin flip in full view of everyone. The questioner invented ten hard general-knowledge questions from their own interests; the contestant answered aloud and usually missed most. Contestants and uninvolved observers then rated both players' general knowledge. Questioners were rated well above average and contestants well below, though every rater had watched chance hand out the advantage of choosing the questions.

Reading the research: An experiment. Role assignment is the independent variable and rated general knowledge the dependent variable, and the coin flip rules out better-informed students having become questioners. The limitation is the setting: a brief staged quiz among students, not an employer judging an employee over months.

Apply it

A classmate arrives late twice in one week. What do you conclude, and what do you skip?

The dispositional attribution comes first and almost automatically: she's disorganized, she doesn't care. The situational explanation gets skipped: a bus that runs every forty minutes, a sibling she drops off first, a teacher who holds the previous room late. Name the pattern: this is the fundamental attribution error, over-weighting character because it is visible and under-weighting circumstances that are not.

Exam tip: concept application items describe a behavior and ask what the observer did. Say which explanation was chosen and which was ignored.

Write it

Identify the research method used in this study and explain what the coin flip rules out.

Show a model response

The researchers used an experiment. Roles were assigned by a coin flip, which is random assignment, so questioners and contestants should not have differed in actual knowledge before the game began. That rules out the explanation that the questioners simply knew more, leaving the role itself as the cause of the ratings gap.

Why it earns the point: part A wants the method named, and the follow-up ties random assignment to the specific alternative it eliminates.

Lesson 4.2 · Unit 4 · CED topic 4.2

Attitude formation and attitude change

The intuitive model is that attitudes drive behavior: you believe, then you act. Half of social psychology is the discovery that the arrow also runs backwards. Get someone to act first (publicly, freely, for no good reason), and the attitude follows the act.

Key terms
  • Attitude: a learned evaluation of a person, object, or idea, carrying belief, feeling, and a tendency to act.
  • Cognitive dissonance: the discomfort of holding an attitude and a behavior that clash; the cheapest repair is usually to move the attitude.
  • Central and peripheral routes: the elaboration likelihood model: motivated, able audiences are moved by argument quality; everyone else by cues such as attractiveness or the source's credibility, meaning expertise plus trustworthiness.
  • Mere exposure effect: repeated contact with a stimulus, with nothing else added, increases liking for it.
  • Foot-in-the-door and door-in-the-face: a small request first makes a large one likelier; an outrageous request first makes a moderate one likelier.
  • Self-fulfilling prophecy: an expectation changes how people act until the expectation comes true.
Study spotlight

Study: Festinger & Carlsmith (1959)

Seventy-one male students spent an hour on deliberately tedious tasks: turning pegs a quarter turn at a time, loading and unloading spools. Each was then asked to tell the next participant in the waiting room that the work had been enjoyable, and was paid either $1 or $20 to say it; a control group said nothing. Interviewed afterwards about how much they had actually enjoyed the tasks, on a scale running from −5 to +5, the $1 group averaged about +1.35 while the $20 group landed near zero.

Reading the research: A laboratory experiment. Payment is the independent variable and rated enjoyment the dependent variable, and the deception required debriefing. Limitations: an all-male undergraduate sample, and $20 in 1959 bought far more than $20 buys now, so the payments were further apart than they look.

Apply it

For a class exercise, a student writes a persuasive essay arguing the opposite of what she believes. No credit, no reward, and she chose the assignment.

Predict that her attitude shifts toward the position she argued. The mechanism is cognitive dissonance under insufficient justification: she argued something she does not believe, and there is no payment or requirement to point at, so the only available explanation for her own behavior is that she partly agrees. Pay her well for the same essay and the shift shrinks: the money explains the act.

Exam tip: dissonance items turn on how weak the external reason is. Big reward, small attitude change.

Write it

Describe what the difference between the two groups' enjoyment ratings indicates about cognitive dissonance.

Show a model response

The $1 group's higher average rating indicates that the students with the weaker external justification changed their attitude more. Twenty dollars was reason enough to explain saying something untrue, so those participants felt little conflict. One dollar was not, so the dissonance between "I said it was fun" and "it was dull" was resolved by deciding the tasks were not so bad after all.

Why it earns the point: part C asks what a statistic indicates, so it reads the direction of the difference as evidence for the mechanism rather than restating the means.

Lesson 4.3 · Unit 4 · CED topic 4.3

The psychology of social situations

Almost everyone believes they would speak up, refuse, or help. These experiments exist because that belief is wrong often enough to be worth measuring. What moves behavior here is rarely the kind of person you are; it is how many people are present, what they appear to think, and whether anyone can tell it was you.

Key terms
  • Normative vs. informational social influence: going along to be accepted, versus going along because you take the group as evidence about reality.
  • Conformity, compliance, obedience: matching a group with no request made, agreeing to a direct request, and following an order from an authority.
  • Group polarization and groupthink: discussion among like-minded people pushes the average view further out; groupthink suppresses dissent to protect harmony.
  • Social facilitation, social loafing, deindividuation: an audience improves easy tasks and worsens hard ones; effort drops when contributions are pooled; anonymity in a crowd loosens self-restraint.
  • Diffusion of responsibility and the bystander effect: the more witnesses, the less any one feels answerable. Darley & Latané (1968) staged an emergency heard over an intercom: participants who believed they were the only listener sought help far more often, and faster, than those who believed several others could hear it too.
  • Prosocial behavior, social traps, superordinate goals: acting to benefit others, with altruism the costly form; situations where each person's rational choice ruins the shared outcome; and a goal neither group can reach alone, which reduces hostility between them.
Study spotlight

Study: Asch (1951, 1956)

A male college student joined what he believed was a group of fellow participants for a vision test: say which of three comparison lines matched a standard. Everyone else in the room was a confederate. On 12 of the 18 trials the confederates answered first and unanimously named an obviously wrong line. Participants went along on roughly a third of those critical trials, and about three-quarters conformed at least once. People judging the same lines alone were almost never wrong.

Reading the research: A laboratory experiment. The confederates' unanimous wrong answers are the independent variable and the number of conforming responses the dependent variable, with the alone condition supplying the baseline error rate. Limitations: the judgment was unambiguous, the participants were male college students, and 1950s conformity norms may not be today's.

Apply it

Your group settles on a project plan you are sure will fail. Everyone nods. You nod.

Sort the two reasons; they need different fixes. If you nodded to avoid being the difficult one, that is normative influence: a concern with acceptance, not truth. If four confident people made you doubt your own reading of the assignment, that is informational influence. Asch's follow-ups point to the fix: unanimity does the work, so one ally who dissents first cuts conformity sharply. Say your objection before the vote.

Exam tip: normative influence answers "what will they think of me"; informational influence answers "what's actually true".

Write it

Explain the extent to which these findings can be generalized, citing specific evidence about the participants.

Show a model response

Generalization is limited. The participants were male college students in the United States in the 1950s, so the results may not describe women, older adults, or cultures with different norms about group harmony. That people conform under social pressure has held up broadly, but the size of the effect should not be carried over to other groups.

Why it earns the point: part E wants participant evidence, so it names who was studied before limiting the claim, and separates the effect from its magnitude.

Lesson 4.4 · Unit 4 · CED topic 4.4

Introduction to personality and its assessment

Personality is the part of you that travels: the pattern of thinking, feeling, and behaving that shows up across situations and over years. Four families of theory explain that pattern in incompatible ways, and each built its own way of measuring it. Learn a theory with its test: the measure is the theory made operational.

Key terms
  • Personality: an individual's characteristic and relatively stable pattern of thinking, feeling, and behaving.
  • The four families: psychodynamic (unconscious conflict), humanistic (growth and the self), social-cognitive (thought, situation, and learning interacting), and trait (measurable dimensions people differ on).
  • Objective tests and personality inventories: standardized questionnaires scored the same way every time; the MMPI is the classic example, built by keeping items that empirically separated groups.
  • Projective tests: ambiguous material the test-taker interprets, on the assumption that inner conflicts leak into the reading: the Rorschach inkblots, the Thematic Apperception Test's story cards. Scoring reliability is the standing objection.
  • Self-report vs. observer report: asking the person, or asking people who know them; observers often predict behavior the person will not admit to.
  • Reliability and validity, and the Barnum effect: a usable measure is consistent and measures what it claims. The Barnum (Forer) effect is our readiness to accept a vague description that would fit almost anyone as personally accurate.
Study spotlight

Study: Forer (1949)

Thirty-nine students completed a personality test and were told each would receive an individual profile based on their answers. A week later every student received the same passage, assembled by the instructor from lines lifted out of an astrology column: statements general enough to fit nearly anyone, such as having unused capacity, or being critical of yourself. Students rated how well the profile described them on a 0-to-5 scale. The mean rating was 4.26.

Reading the research: A classroom demonstration experiment with a single condition: everyone received the identical profile, and rated accuracy is the dependent variable. The limitations are structural, no control group received a genuinely individualized profile for comparison, and the participants were the instructor's own students, who had reason to agree.

Apply it

An app charges $30 for a personality report. Yours says you are sociable but value time alone, and that you set high standards for yourself. It feels uncannily right.

That feeling is the Barnum effect: the statements are two-sided or near-universal, so almost any reader confirms them. Ask what a real inventory would show. Reliability: the same person scores about the same weeks later, and items within a scale hang together. Validity: scores predict something independent (observer ratings, job performance) rather than only feeling accurate to the person who paid.

Exam tip: "it described me perfectly" is evidence about the reader, not the test. Validity needs an outside criterion.

Write it

Identify one design flaw in this study and state what a control group would have to receive to fix it.

Show a model response

The flaw is that there is no comparison condition: every student received the generic profile, so a mean of 4.26 cannot be measured against anything. A control group would have to receive a genuinely individualized profile built from their own answers. If both were rated about equally accurate, the ratings would be tracking the reader's willingness to agree rather than the content of the profile.

Why it earns the point: it names the missing feature and specifies what the control condition receives, which is what a scientific investigation item asks for.

Lesson 4.5 · Unit 4 · CED topic 4.5

Psychoanalytic and humanistic theories of personality

These two families disagree about what a person fundamentally is. Freud's picture is a mind at war with itself, managing impulses it cannot admit to having. The humanists answer that people are growing toward something, and that what blocks growth is usually the conditions other people attach to their approval. Both are influential; neither is easy to test.

Key terms
  • The unconscious; id, ego, superego: thoughts outside awareness that still shape behavior; the id demands gratification, the superego moralizes, and the ego negotiates between them and reality.
  • Defense mechanisms: the ego's unconscious distortions: repression, denial, projection (attributing your impulse to someone else), displacement (redirecting it at a safer target), rationalization, regression, reaction formation (acting the opposite), and sublimation (channelling it into accepted work).
  • Psychosexual stages and fixation: oral, anal, phallic, latency, genital; unresolved conflict at a stage was said to mark personality lastingly.
  • The neo-Freudians: Adler's inferiority complex and the drive to compensate; Horney's basic anxiety from childhood insecurity, and her rejection of Freud's account of women; Jung's collective unconscious of shared archetypes.
  • Maslow's hierarchy and self-actualization: physiological, safety, belonging, and esteem needs, with self-actualization at the top.
  • Rogers's terms: unconditional positive regard; congruence, the match between self-concept and ideal self; and conditions of worth, the strings attached to approval that make people disown parts of themselves.
Study spotlight

Study: Maslow (1954)

Maslow wanted to describe psychological health rather than illness, so he built a sample by judgment: living acquaintances he considered to be fulfilling their potential, plus historical figures chosen the same way. He read biographies, letters, and other writings, and interviewed those he could. From that material he drew a list of shared characteristics: problem-centred rather than self-centred, comfortable with solitude, resistant to social pressure, and given to intense peak experiences. He reported no sample size of the kind a modern paper states.

Reading the research: Qualitative case-study analysis, not an experiment. The sample was chosen by the researcher against criteria the researcher defined, which is circular: people picked for being self-actualized shared the qualities he used to pick them. It is rich in description, very low in objectivity and generalizability, and as stated essentially unfalsifiable, no result could have contradicted it.

Apply it

Dev is cut from the team. He tells everyone the coach is the one who feels threatened, and snaps at his younger brother that evening.

A Freudian reads two defenses: projection, attributing his own threatened feeling to the coach, and displacement, redirecting anger onto a safer target. Rogers describes the same evening without an unconscious. Dev's self-concept includes being an athlete, and being cut puts his real experience out of step with that self: incongruence. If the approval around him carries conditions of worth, admitting the hurt feels unsafe, so it comes out sideways. The remedy is unconditional positive regard: being valued while he says the true thing.

Exam tip: projection and displacement are the pair most often confused. Projection moves the feeling onto someone else; displacement moves the target.

Write it

Identify the research method Maslow used, and explain one reason it cannot test his theory.

Show a model response

Maslow used case studies: qualitative analysis of biographical material about people he selected himself. It cannot test the theory because he chose the cases using the very trait he wanted to describe, so the conclusion was guaranteed by the selection. With no comparison group judged not to be self-actualizing, no result could have shown the description wrong.

Why it earns the point: it names the method and identifies the circularity, rather than only saying the sample was small.

Lesson 4.6 · Unit 4 · CED topic 4.6

Social-cognitive and trait theories of personality

The two modern families are the testable ones. Social-cognitive theory explains personality as the loop between what you think, what you do, and the situations your behavior keeps putting you in. Trait theory skips explanation and measures: find the dimensions people reliably differ on, and see what those scores predict.

Key terms
  • Reciprocal determinism: Bandura's claim that personal factors, behavior, and environment each influence the other two; a sociable person chooses parties, and parties make them more sociable.
  • Self-efficacy and self-esteem: your belief that you can carry out a specific task, versus your overall sense of your own worth. Self-efficacy is task-specific and predicts persistence.
  • Locus of control and personal control: an internal locus credits outcomes to your own effort, an external locus to luck or powerful others.
  • Trait and factor analysis: a stable pattern of behavior, and the statistical method that finds which rating items cluster together. Eysenck reduced personality to a few broad dimensions, among them introversion-extraversion.
  • The Big Five: openness, conscientiousness, extraversion, agreeableness, and neuroticism: the five factors that recur across languages and cultures.
  • The person-situation controversy: the long argument over whether traits or situations better predict behavior. The settlement: traits predict behavior averaged over many occasions; the situation usually wins on any single occasion.
Study spotlight

Study: Mischel and colleagues (1972 onward); Watts, Duncan & Quan (2018)

Mischel's team left a preschooler alone with one marshmallow and a promise: wait until the experimenter returns and you get two. Waiting times varied enormously, and later follow-ups of these few dozen children, nearly all attending a single university's nursery school, reported links between waiting longer and better adolescent outcomes. Watts, Duncan and Quan repeated the design with more than 900 children drawn from a far more socioeconomically diverse sample. The association with later achievement was roughly half the size, and it shrank further once family background and early cognitive ability were statistically controlled.

Reading the research: The original is a small-sample experiment with correlational follow-up: waiting time was measured, not assigned, so the link to later outcomes was never causal. The replication adds statistical control and a broader sample. The limitation of both is that a single delay task on a single afternoon is a narrow operational definition of self-control.

Apply it

A hiring manager wants to give applicants a Big Five inventory and use the conscientiousness score to predict who will perform best next Tuesday.

Tell her what the score can and cannot do. Conscientiousness does predict job performance averaged over months, because a small tendency repeated across hundreds of occasions accumulates. It predicts one Tuesday poorly: that day is dominated by the situation; workload, deadline, who else is in the room, whether the person slept. Predicting a single act from a trait score is the error the person-situation controversy settled.

Exam tip: when an item asks what a trait predicts, the answer is aggregated behavior, not a specific act.

Write it

Explain what the replication's smaller association indicates about the original conclusion.

Show a model response

It indicates the original conclusion was overstated rather than wrong. Waiting longer still goes with somewhat better later outcomes, but the relationship is much weaker in a larger, more diverse sample, and part of what looked like an effect of self-control was really family background and early cognitive ability doing the work. Because both studies measured waiting rather than assigning it, neither shows that teaching a child to wait would raise achievement.

Why it earns the point: it reports the direction and size of the change, names the controlled variables, and refuses the causal claim.

Lesson 4.7 · Unit 4 · CED topic 4.7

Motivation

Motivation theories fall into two groups, and the exam wants you to tell them apart. Push theories say an internal deficit drives you until it is corrected. Pull theories say the reward out in the world draws you toward it. Neither alone explains why someone with a full stomach eats dessert, or why a paid hobby stops being fun.

Key terms
  • Drive-reduction theory and homeostasis: a physiological need creates a drive; behavior that restores the body's set point reduces it. Incentives are the pull side: external rewards that motivate regardless of need.
  • Arousal theory and the Yerkes-Dodson law: people act to keep arousal at a comfortable level, and performance peaks at moderate arousal: lower for hard or new tasks, higher for simple well-practised ones. Sensation seeking is a stable preference for a high level.
  • Intrinsic and extrinsic motivation: doing something for its own sake, versus for a reward or to avoid a punishment.
  • Overjustification effect: paying someone for something they already enjoyed can replace the internal reason with the external one, so the behavior drops when payment stops.
  • Need hierarchies and self-determination theory: Maslow's hierarchy orders needs from physiological and safety up to self-actualization; self-determination theory holds that motivation is sustained by three needs, autonomy (choice), competence (getting better), and relatedness (connection to others).
  • Motivational conflicts and hunger: approach-approach (two good options), avoidance-avoidance (two bad ones), approach-avoidance (one option with both). Hunger is regulated by the hypothalamus reading signals including ghrelin, which rises before meals, and leptin, from fat cells. Achievement motivation is the drive toward challenging standards.
Study spotlight

Study: Deci (1971)

About two dozen college students worked on an interesting spatial puzzle across three sessions. Half were paid per puzzle solved during the middle session only; the rest were never paid. In each session the experimenter left the room for an eight-minute break, telling students they could do whatever they liked; magazines lay on the table. The measure was how many of those unmonitored seconds each student spent on the puzzle anyway. In the final session, after payment had stopped, the previously paid students spent less of that free-choice period on the puzzle than the never-paid students did.

Reading the research: An experiment. Payment is the independent variable; the dependent variable is seconds spent on the puzzle during the free-choice period: an elegant behavioral operational definition of intrinsic motivation, since nothing external rewards the choice. The limitations are small samples and a very short time frame.

Apply it

A parent pays a twelve-year-old $5 for every book finished. Reading climbs for two months. Then the payments stop.

Predict that reading falls below where it started. This is the overjustification effect: the child now explains their own reading by the money, so when the money goes the reason goes with it. Redesign it with self-determination theory. Autonomy: the child picks the books. Competence: track pages this month against last, so progress is visible. Relatedness: someone reads the same book and talks about it. None of the three can be withdrawn the way a payment can.

Exam tip: reinforcement raises a behavior while it is delivered; overjustification is what happens to an already enjoyed behavior after it stops.

Write it

State the operational definition of intrinsic motivation used in this study.

Show a model response

Intrinsic motivation was operationalized as the number of seconds a student spent working on the puzzle during an eight-minute free-choice period when the experimenter had left the room, no payment was available, and magazines offered an alternative. It is a behavioral count under conditions where no external reward could explain the choice, not a rating of how interesting the puzzle felt.

Why it earns the point: part B wants the procedure behind the number; what was counted, for how long, and under what conditions.

Lesson 4.8 · Unit 4 · CED topic 4.8

Emotion

Every theory here answers one question in a different order: when your heart pounds and you feel afraid, which came first? Learn them as sequences rather than names: that is how the exam tests them, with a scenario and a question about what has to happen first.

Key terms
  • James-Lange theory: the body reacts first and the emotion is your reading of that reaction: you are afraid because you are trembling.
  • Cannon-Bard theory: the thalamus routes the signal so that arousal and the felt emotion happen at the same time, neither causing the other.
  • Schachter-Singer two-factor theory: emotion requires physical arousal plus a cognitive label for it, taken from the situation. Same racing heart, different label, different emotion.
  • Lazarus's cognitive appraisal: an appraisal of what the event means for you, often fast and unconscious, determines which emotion follows. Broaden-and-build theory adds that positive emotions widen attention and thinking, accumulating durable resources over time.
  • Facial feedback hypothesis: expressions feed back into what is felt, so holding a smile should nudge mood upward. Treat it cautiously: a large multi-laboratory replication failed to reproduce the classic finding.
  • Expression and biology: several basic emotions are recognized from faces across cultures, while display rules govern which may be shown where. The amygdala tags threat fast. A polygraph measures autonomic arousal, not lying, which is why an innocent person afraid of being disbelieved can fail one.
Study spotlight

Study: Schachter & Singer (1962)

A total of 184 male college students were told they were receiving a vitamin injection in a study of vision. Most were given adrenaline, which raises heart rate and causes trembling; some received an inert placebo. Of those given adrenaline, some were told the true side effects, some were told nothing, and some were misinformed. Each then waited with a confederate who behaved either euphorically or angrily. Participants with no explanation for their pounding hearts reported and displayed emotions matching the confederate; those who could blame the injection did not.

Reading the research: A factorial experiment: what participants were told and how the confederate behaved are the two independent variables, self-reported and observer-rated emotion the dependent variables. It relied on deception, so debriefing was required. The limitation is strength of evidence: the effects were weaker than the theory's fame suggests, and later replications have been mixed.

Apply it

People who meet a stranger halfway across a swaying footbridge over a gorge tend to rate that stranger as more attractive than people who meet the same stranger on level ground.

Two-factor theory explains it. The bridge produces the arousal (pounding heart, shallow breathing), and the person labels it using the most available feature of the situation, the attractive stranger in front of them. The label is wrong, and that is the point: this is misattribution of arousal. Note what the explanation requires: the arousal has to come first and be unexplained.

Exam tip: if a scenario gives an obvious physical cause for arousal and the person still attributes it to emotion, the item is testing two-factor theory.

Write it

Identify one ethical guideline these researchers were required to apply, and explain why the procedure made it necessary.

Show a model response

Debriefing. Participants were deceived twice (told the injection was a vitamin, and in one condition misinformed about its effects), and the design depends on the arousal being unexplained, so the study could not have worked otherwise. Deception is permitted only when the question requires it, so afterwards participants must be told the true purpose and what they actually received, with the option to withdraw their data.

Why it earns the point: part D wants a guideline named and tied to a specific feature of the procedure, not a general statement that ethics matter.

Lesson 4.9 · Unit 4 · Research methods

Sampling and generalizability

A finding is about the people it was measured on. Extending it to everyone else is a separate claim, and the only thing that supports it is how the sample was drawn. Social psychology has the sharpest version of this problem, because the behavior it studies (conformity, attribution, what counts as an insult) is exactly the kind that varies between cultures.

Key terms
  • Population and sample: everyone the conclusion is meant to describe, and the subset actually studied.
  • Random selection: every member of the population has an equal chance of being chosen. It licenses generalization, and is not the same as random assignment, which licenses causal claims.
  • Convenience and stratified samples: taking whoever is easy to reach, versus sampling within subgroups so their proportions match the population.
  • Sampling bias and volunteer bias: a sample systematically unlike the population; self-selection is one route, since people who answer a survey about stress may differ from those who ignore it.
  • WEIRD samples: participants from Western, educated, industrialized, rich, democratic societies, who are heavily overrepresented in published psychology.
  • External and ecological validity: whether results extend beyond this study's people and setting, and whether the task resembles real life. Cross-cultural replication is how either one is actually tested.
Study spotlight

Study: Henrich, Heine & Norenzayan (2010)

Rather than run a new experiment, these researchers examined who had been studied in leading behavioral science journals. Roughly 96% of sampled participants came from Western industrialized countries, and about 68% from the United States alone: countries holding around 12% of the world's population. They then compared results across societies on tasks assumed to be universal. On the Müller-Lyer illusion and on economic bargaining games, Western undergraduates sat at one end of the range rather than in the middle, with some small-scale societies showing little of the illusion at all.

Reading the research: A review of published methods sections combined with cross-cultural comparisons, not an experiment, so nothing here is manipulated and no causal claim is made. The limitation is scope: it describes who gets published, which shows that many findings are untested elsewhere, not that any particular finding is false.

Apply it

A study of 60 undergraduates at one American university, mostly first-year psychology students earning course credit, about 70% women, finds that self-testing beats rereading. Write part E.

Generalization is limited, and the useful answer names which features matter and why. All participants were young adults at a single American university, so the results may not extend to older learners, to people who did not attend college, or to educational cultures that teach studying differently. They volunteered for course credit, so volunteer bias is possible. The sex imbalance is the weaker objection: with no strong reason to expect retrieval practice to work differently for women and men, it limits the claim less than the others do.

Exam tip: part E is scored on participant evidence. "The sample was small" alone rarely earns it: name a characteristic and say what it would change.

Write it

Identify the sampling method used in that undergraduate study and explain one bias it introduces.

Show a model response

It is a convenience sample: the researchers recruited students easy to reach on their own campus rather than selecting randomly from a defined population. One bias this introduces is sampling bias toward people with strong recent study habits, since undergraduates are selected for academic skill and practice memorizing constantly. A technique that helps them may do less for learners who study rarely, so the effect could be overstated for the wider population.

Why it earns the point: it names the method and traces one specific route from the recruiting procedure to a distorted result.

Lesson 4.10 · Unit 4 · Research methods

Inferential statistics: p-values and what "significant" means

Descriptive statistics summarize the sample you have. Inferential statistics answer a harder question: could a difference this large have arisen by chance, in a world where the two groups are really the same?

Key terms
  • Null hypothesis: the claim that there is no real difference or relationship in the population; the analysis tries to make it look implausible.
  • p-value: the probability of a result at least this extreme if the null hypothesis were true. It is not the probability the null hypothesis is true, nor the probability the finding will replicate.
  • Statistical significance: by convention a result is significant when \(p \lt .05\), an agreed-upon cutoff rather than a fact about nature. It supports inferring the effect in the population the sample came from; whether it holds elsewhere is a sampling question.
  • Type I and Type II error: a Type I error rejects a true null hypothesis: a false alarm. A Type II error misses a real effect. Loosening the cutoff trades one for the other.
  • Sample size and power: power is the chance of detecting a real effect when there is one. Larger samples raise power, which is why a big enough study can make a trivial difference significant.
  • Effect size: how large the difference or relationship is, reported as d or r. Significance says an effect probably exists; only effect size says whether it matters.
Study spotlight

Study: Open Science Collaboration (2015)

Some 270 researchers collaborated on the Reproducibility Project: Psychology, repeating 100 studies published in three well-regarded journals, using the original materials and consulting the original authors wherever possible. Of the original studies, 97% had reported statistically significant results. Of the replication attempts, only about 36% reached significance, and their effects averaged roughly half the size of the originals.

Reading the research: A coordinated set of direct replications, not a new experiment. The dependent variables are whether each replication reached significance and how large its effect was. The limitation is interpretive: a replication can fail because the original was a false positive, but also because of an unknown moderator or a procedure that did not transfer.

Apply it

A company advertises its study app with this results table. Read it carefully.

GroupnMean test scoreSD
Used the app48075.610.0
Control48074.210.0

Reported with the table: mean difference 1.4 points, p = .03, d = 0.14.

What p = .03 tells you: if the app truly made no difference, a gap this big or bigger would turn up about 3 times in 100 by chance, so the result clears the usual cutoff. What it does not tell you: how big the effect is, whether it matters, or the probability the app works. Effect size settles that: d = 0.14 is small, and it cleared the cutoff largely because 960 students were tested. A Type I error is live too: this may be a false alarm.

Exam tip: data analysis items love a significant result with a small effect size. Answer both: significant, and small.

Write it

Describe what the replication rate of about 36% indicates about the published literature.

Show a model response

It indicates that a published significant result is a weaker guarantee than readers assumed: nearly all the originals were significant, only about a third of the repeats were, and the repeated effects were about half as large. That fits a literature in which effects get published when they clear \(p \lt .05\) and stay unpublished when they do not, so published estimates are inflated. It does not establish that two-thirds of those findings are false.

Why it earns the point: part C asks what a statistic indicates, so it reads the gap between the two rates and stops short of overclaiming.

Unit 4 review · 10 multiple-choice

Unit 4 review: Social Psychology and Personality

Ten questions across all ten lessons in the exam's three styles (concept application, data analysis, and scientific investigation), so answer each one before you click, then read all four rationales, not just the one for the option you chose.

Multiple choice

  1. Ravi's teammate Nia misses two group meetings, and Ravi concludes she is unreliable and does not care about the project. When Ravi himself misses the next meeting, he explains that his shift ran late and the bus never came. Which of the following best describes Ravi's pattern of explanation?

    Correct. Ravi reaches for Nia's character (unreliable, uncaring), and skips the circumstances, which is the fundamental attribution error; when the target is himself he reaches for the bus and the shift instead. That reversal by target is exactly what the actor-observer bias names.

    The directions are backwards: Ravi credits Nia's disposition, not her situation. The just-world hypothesis is also a different idea: the assumption that people get the outcomes they deserve, which would show up as "she must have done something to deserve it," not as an explanation of lateness.

    False consensus is overestimating how widely your own view or behavior is shared, as in assuming most people would also skip the meeting. Nothing in the scenario tells you what Ravi believes about other classmates, so this is a real bias that does not answer the question.

    A halo effect runs from one positive impression outward, letting a single good quality inflate unrelated judgments. Ravi is not generalizing from one trait across many; he is explaining one behavior, and the two explanations differ depending on whose behavior it is.

  2. A club needs volunteers to hand out flyers for a cause its members feel lukewarm about. Half the members are offered a $20 gift card to do it; the other half are asked as a favour and receive nothing. Everyone hands out flyers. A week later, which group is likelier to report warmer attitudes toward the cause, and why?

    Reinforcement raises the rate of a behavior while it is delivered; it does not reliably move the private attitude behind it. In dissonance research the payment works the other way, because a large reward supplies a ready explanation for the act and leaves no discomfort to resolve.

    Correct. Dissonance is largest under insufficient justification: the unpaid members advocated for something they barely believed, with no money or requirement to point at, so the cheapest repair is to shift the attitude toward the behavior. This is the Festinger and Carlsmith pattern, in which the $1 group reported more enjoyment than the $20 group.

    Mere exposure is increased liking from repeated contact with a stimulus itself, not from a reward attached to it, and a gift card given for handing out flyers is an incentive rather than a peripheral persuasion cue. The distractor mixes two real concepts and applies neither correctly.

    Attitude change through the central route is one path, but it is not the only one: the elaboration likelihood model also allows peripheral change, and dissonance changes attitudes with no persuasive message at all. Here the members' own behavior, not an argument, is what does the work.

  3. Described data

    A research team ran a line-judgment task. On each trial the participant said aloud which of three comparison lines matched a standard line; the correct answer was obvious. Confederates seated with the participant answered first. The table reports the percentage of critical trials on which participants gave the majority's wrong answer, by condition.

    ConditionPercent of critical trials conformed
    Judging alone, no group present1%
    Unanimous majority of 332%
    Unanimous majority of 736%
    Majority of 7, one confederate answers correctly7%

    Which conclusion is best supported by the data in the table?

    The table contradicts this. Going from three confederates to seven, more than doubling the majority, moved conformity only from 32% to 36%, which is the finding that majority size stops mattering quickly once a few people agree.

    The alone condition rules this out: with no group present, participants were wrong on about 1% of trials, so they could see perfectly well which line matched. That baseline is what turns the other rows into evidence about social pressure rather than about perception.

    Correct. Compare the rows the data actually let you compare: 36% minus 32% is about 4 percentage points for four extra people, while 36% minus 7% is about 29 points for one confederate who breaks the unanimity. The comparison also separates the mechanisms: a single ally mostly relieves normative pressure, the concern about being the odd one out.

    Read the column heading. These are percentages of trials, not of participants, and the two are very different claims: a third of trials can come from most people conforming occasionally rather than from a third of people conforming always.

  4. Described study

    A researcher posts a 40-item online personality inventory and recruits 60 students to complete it. A week later every student receives a paragraph described as "your individual profile," and rates on a 1-to-5 scale how accurately it describes them. In fact all 60 students receive the identical paragraph, made of statements general enough to fit almost anyone. The mean accuracy rating is 4.2. The researcher concludes that the inventory is a valid measure of personality.

    Which change to the design would most directly test the researcher's conclusion?

    A larger sample narrows the confidence interval around 4.2, but it cannot tell you what the 4.2 means, because there is still nothing to compare it against. Precision is not the problem here; the missing comparison condition is.

    Stable ratings across two weeks would show only that students agree with the same vague paragraph twice. Reliability is consistency of measurement, and a measure can be perfectly consistent while measuring nothing the researcher claims: reliability is necessary for validity but never sufficient.

    Inter-rater reliability matters when human judges score open-ended material, as with a projective test. Here the students score themselves on a fixed scale, so there are no raters to agree, and agreement would still say nothing about whether the inventory captures personality.

    Correct. With every student receiving the same generic text, a mean of 4.2 is fully explained by the Barnum effect: near-universal statements feel personal. A control group given genuinely individualized profiles supplies the comparison: only if the individualized profiles are rated clearly more accurate is the inventory adding information.

  5. Marcus is passed over for a solo. He tells his friends the choir director is jealous of his voice, and that evening he snaps at his younger sister over something trivial. A psychoanalytic theorist would name these two responses, in order, as

    The order is reversed. Telling everyone the director is the jealous one comes first and moves a feeling onto another person, which is projection; snapping at a sister comes second and moves a target, which is displacement.

    Correct. Projection attributes your own unacceptable feeling (here, envy) to someone else. Displacement keeps the feeling and swaps the target for a safer one, which is why the anger aimed at the director lands on a younger sibling instead. Remember the pair as: projection moves the feeling, displacement moves the target.

    Reaction formation would have Marcus acting conspicuously delighted for whoever got the solo, expressing the opposite of what he feels. Rationalization would have him producing an acceptable reason for the outcome, such as deciding he never wanted the part: neither matches what he actually does.

    Rationalization supplies a comfortable justification rather than accusing someone else of the feeling. Regression would mean returning to behavior from an earlier developmental stage, such as a tantrum, and snapping once at a sibling is ordinary displaced irritation, not a return to childhood functioning.

  6. Talia is confident she can work through a hard calculus problem set if she keeps at it, and she persists for an hour when a problem resists her. She also believes she could never give a decent speech, and she avoids the one presentation in the course. Which construct does this pattern best illustrate?

    Locus of control is a broad belief about what produces outcomes in general, your own effort versus luck and powerful others, so it would not flip between two courses in the same semester. The scenario gives you a belief about capability in a specific domain, not about where outcomes come from.

    Self-esteem is a global evaluation of your own worth, which is why it changes slowly and does not vary by subject. Talia's confidence is high in one task and low in another, and that variation is the clue that the item is testing a task-specific construct.

    Correct. Self-efficacy is Bandura's term for the belief that you can carry out a particular task, and it is domain-specific by definition: high for calculus, low for public speaking, in the same person. It also predicts exactly what the scenario reports: persistence in the high-efficacy task and avoidance in the low-efficacy one.

    Reciprocal determinism is the three-way loop among personal factors, behavior, and environment. It would apply if the scenario showed the avoidance changing her environment, which then fed back onto the belief; as written, you are given the belief and its immediate behavioral consequence.

  7. Described data

    Students worked on an interesting spatial puzzle in three sessions. In each session the experimenter left the room for an 8-minute (480-second) break and said the students could do whatever they liked; magazines lay on the table. In session 2 only, students in one group earned $1 per puzzle solved. The table gives each group's mean number of seconds spent on the puzzle during the unmonitored break.

    GroupSession 1Session 2Session 3
    Paid in session 2205 s290 s155 s
    Never paid208 s215 s220 s

    Which statement is best supported by the table?

    Correct. Track the paid row against its own session 1 baseline: 205 to 290 while the dollar was on offer, then down to 155, fifty seconds below where it started, after payment stopped. The never-paid row barely moves, which rules out simple boredom with the puzzle and leaves the reward as the difference between the groups.

    The session 2 mean is indeed the table's highest, but the claim is about what happens afterwards, and session 3 answers it: 155 seconds is the table's lowest value. A reward can raise behavior while it is delivered and still leave the behavior lower than baseline once it is withdrawn.

    Their means rise slightly across the three sessions, from 208 to 215 to 220 seconds. That flat-to-upward trend is what makes this group a useful comparison: whatever the paid group's session 3 drop reflects, it is not the puzzle becoming dull with repetition.

    Negative reinforcement means a behavior increases because it removes something aversive, and nothing aversive is removed here. Payment is a positive reinforcer during session 2; the interesting result is the overjustification effect that follows it, which is a change in the reason for the behavior rather than a schedule of reinforcement.

  8. Ana finishes a hard training run and meets a stranger at the finish line while her heart is still pounding and her breathing has not settled. She later reports finding the stranger unusually attractive. Which explanation follows Schachter and Singer's two-factor theory?

    This is James-Lange, in which the emotion just is your reading of a specific bodily pattern and no cognitive label from the situation is needed. Two-factor theory was built precisely because the same racing heart can become fear, anger, or attraction depending on what is around you.

    This is Cannon-Bard: arousal and emotional experience occur simultaneously, with the thalamus routing the signal both ways and neither one producing the other. It is a real theory but it cannot explain why the stranger, rather than the run, gets the credit.

    The facial feedback hypothesis says expressions feed back into what is felt, and it is worth treating cautiously since a large multi-laboratory replication failed to reproduce the classic result. It also does not fit: the scenario gives you exertion and a pounding heart, not a held smile.

    Correct. Two-factor theory requires both ingredients, physical arousal plus a cognitive label drawn from the situation, and the label is wrong here, which is what misattribution of arousal means. Notice the condition the explanation depends on: the arousal has to be unexplained, so if Ana consciously credited the run, the effect should disappear.

  9. Described study

    Researchers post a survey link in an online forum devoted to sleep advice. Over two weeks 1,200 forum members complete it, answering questions about their electronics use and their sleep. Among respondents who report keeping a screen curfew in the hour before bed, 68% rate their sleep quality as good or very good, compared with 41% of respondents with no curfew. The researchers conclude that most adults would sleep better if they adopted a screen curfew.

    Which of the following is the most serious threat to the researchers' conclusion?

    Twelve hundred respondents is a large sample, and a 27-percentage-point gap is not a subtle effect, so power is not the issue. Sample size and sample representativeness are different problems: a huge sample drawn from the wrong population is still the wrong population.

    Correct. People who join a forum about sleep advice are already unusual in their sleep habits and their interest in fixing them, and those who answered are volunteers within that group, so this is volunteer bias layered on a convenience sample. Part E of an AAQ is scored on this kind of participant evidence: name who was recruited and where, then say what it does to the claim.

    Self-report has real weaknesses, social desirability and response sets among them, but it is standard, usable evidence when the construct is a private experience such as perceived sleep quality. An answer that rejects a whole class of measurement is too strong to be the best objection here.

    This confuses the two randoms. Random assignment licenses causal claims, and its absence does mean the curfew may not be what produced better sleep; random selection is what licenses generalizing to a population. The conclusion quoted is a generalization, so the sampling term is the one that applies.

  10. Described data

    Students were randomly assigned to use a vocabulary app or to a control activity for four weeks, then took a 100-point vocabulary test. Reported with the table: the mean difference is 1.4 points, p = .02, and Cohen's d = 0.12.

    GroupnMean test scoreSD
    Used the app60071.812.0
    Control activity60070.412.0

    Which statement is best supported by these results?

    Significance and size are separate questions, and this is the confusion data analysis items are built to catch. A p-value says a difference this large is unlikely if the app truly did nothing; only the effect size says whether the difference is worth anything, and d = 0.12 is small by any convention.

    A p-value is the probability of a result at least this extreme if the null hypothesis were true, not the probability that the null hypothesis is true, and not the probability the finding will replicate. Reversing the conditional is the single most common error with p-values.

    Correct. Both readings are in the numbers: 71.8 minus 70.4 is a 1.4-point gap, p = .02 clears the conventional .05 cutoff, and 1.4 divided by the standard deviation of 12.0 gives d ≈ 0.12. Power is why both are true at once, with 600 students per group, even a trivial difference becomes detectable, which is exactly why effect size is reported alongside.

    A Type II error is missing a real effect: failing to reject a null hypothesis that is false. These researchers detected an effect, so if they erred it would be a Type I error, a false alarm; and a small effect size is not itself evidence of any error.

Lesson 5.1 · Unit 5 · CED topic 5.1

Introduction to health psychology: stress and coping

Stress is not the exam. Stress is what your body does about the exam: an appraisal that a demand outruns your resources, then a response built for predators and aimed at a calendar. A system tuned for minutes must run for months, and the bill arrives in the immune system.

Key terms
  • Stressor and stress: the stressor is the event; stress is appraising it as threatening and responding.
  • Eustress and distress: arousal that sharpens performance, versus arousal that swamps it.
  • General adaptation syndrome: Selye's three stages: alarm, resistance at a costly plateau, then exhaustion, when illness risk rises.
  • HPA axis and cortisol: hypothalamus to pituitary to adrenal cortex, releasing cortisol. That is the slow arm; fight-flight-freeze is the fast one.
  • Tend-and-befriend: an alternative threat response: protect dependents, seek contact. Social support buffers the damage.
  • Coping: problem-focused coping changes the stressor, emotion-focused coping changes your reaction. Psychoneuroimmunology links those states to immune function; burnout is chronic unrelieved demand.
Study spotlight

Study: Cohen, Tyrrell & Smith (1991)

Three hundred ninety-four healthy adults answered questionnaires combining recent stressful life events, perceived stress, and negative mood into one psychological stress index. Each then received nasal drops containing a common respiratory virus and was quarantined while staff recorded symptoms and laboratory tests confirmed infection. Rates of infection, and of colds verified by those tests rather than by self-report, climbed steadily across the stress index, from roughly a quarter of the least-stressed participants to nearly half of the most-stressed.

Reading the research: A prospective study bolting a measured variable onto an experimental one. Exposure was controlled, but the stress index was measured rather than assigned, so the stress-illness link stays correlational. The dependent variable was a clinically verified cold. The limitation is selection: adults willing to be quarantined are a cross-section of nobody.

Apply it

Dev says exam week wrecks him: two sleepless nights, four days of grinding, then a sore throat the morning finals end.

That is the general adaptation syndrome in order. The sleepless nights are alarm; the grinding days are resistance, cortisol holding a costly plateau; the sore throat on the first free morning is exhaustion, when suppressed immune function meets a virus that was there all along. Problem-focused: a fixed revision schedule. Emotion-focused: a run, which changes nothing about the exam and much about him.

Exam tip: application items live on this split. Ask whether the strategy touches the stressor or the reaction.

Write it

Explain the extent to which this study's findings can be generalized, citing specific evidence about the participants.

Show a model response

Generalization is real but limited. The participants were 394 healthy adults who volunteered to be given a virus and quarantined, so they were screened for health and self-selected for unusual willingness. The sample is large enough that the graded pattern is unlikely to be a fluke, but it should not be stretched to children or to serious infections.

Why it earns the point: part E wants participant evidence, so it names who was screened in and what that costs the claim.

Lesson 5.2 · Unit 5 · CED topic 5.2

Positive psychology: strengths, well-being, and what lasts

For most of its history psychology studied what goes wrong. Positive psychology asks the other question, what makes a life go well, and asks it with the same tools: random assignment, control conditions, follow-up measurement. The surprise is how little of what people chase actually moves the number.

Key terms
  • Subjective well-being: self-reported life satisfaction plus the balance of positive to negative feeling. The field's main outcome measure.
  • Gratitude and resilience: noticing and crediting what has gone well; adapting successfully to adversity rather than avoiding it.
  • Grit: sustained passion and perseverance toward long-term goals, which predicts outcomes partly independently of ability.
  • Flow: absorbed, self-forgetting engagement in a task whose difficulty matches your skill.
  • Signature strengths and virtues: the character strengths a person uses most naturally, grouped under virtues such as courage, humanity, and justice.
  • Adaptation level and the hedonic treadmill: you judge new experiences against your recent level, so gains fade toward baseline. Post-traumatic growth is the opposite case: positive change reported after struggle.
Study spotlight

Study: Seligman, Steen, Park & Peterson (2005)

Five hundred seventy-seven adults who found the study online were randomly assigned to one of five one-week exercises or to a placebo control that wrote about early memories. Happiness and depressive symptoms were measured before, immediately after, and at one, three, and six months. Two exercises outperformed the placebo long past the week they were done: writing down three good things each day with what caused them, and using a signature strength in a new way. Both raised happiness and lowered depressive symptoms for up to six months.

Reading the research: A randomized controlled trial with a placebo condition, so expectancy is held roughly constant across groups. The independent variable is the assigned exercise; the dependent variables are scores on self-report happiness and depression scales. Two limitations: an internet sample already interested in becoming happier, and heavy attrition, since the people who keep answering at six months are probably the ones it worked for.

Apply it

Design a two-week gratitude intervention for your class. Name the control condition, the operational definition, and the ethical steps.

Randomly assign students to write three good things each night with their causes, or to a control that logs three things that happened, with no causes: that holds writing time constant. Operationalize well-being as the mean of a five-item life-satisfaction scale completed on day 1 and day 15. Ethically: parental consent plus assent for minors, a free choice to stop, no grade attached, confidential responses, and the exercise offered to the control group afterward.

Exam tip: a control condition that does nothing is weak. Match everything except the active ingredient.

Write it

Identify the research method used in this study and state the operational definition of well-being.

Show a model response

The researchers used an experiment: a randomized controlled trial, since participants were randomly assigned to an exercise or to a placebo writing task. Well-being was operationalized as participants' scores on self-report happiness and depressive-symptom questionnaires, completed before the exercise and again at one, three, and six months.

Why it earns the point: part A names the method rather than retelling the procedure, and part B gives the measuring instrument and the schedule, not a definition of happiness.

Lesson 5.3 · Unit 5 · CED topic 5.3

Explaining and classifying psychological disorders

Deciding that a pattern of thought, feeling, or behavior is a disorder is a judgment, and psychology tries to make it a disciplined one. No single sign settles it: grief is intense distress that is not a disorder, and a person can be seriously unwell without seeming odd to anyone. A diagnosis names a pattern, not a person: a working summary that guides treatment and can be revised.

Key terms
  • DSM-5-TR and ICD: the American Psychiatric Association's manual of diagnostic criteria, and the World Health Organization's international classification.
  • Dysfunction, distress, deviation, danger: the criteria clinicians weigh: interference with daily functioning, personal suffering, departure from cultural norms, risk of harm. Any one alone is a poor test.
  • Medical and biopsychosocial models: disorder as illness with causes and treatments, versus disorder as biological, psychological, and sociocultural factors together.
  • Diathesis-stress model: a vulnerability, inherited or acquired, becomes disorder only when enough stress activates it.
  • Comorbidity, eclectic approach, stigma: two or more diagnoses at once; a clinician drawing on several theories; the labeling effect that changes how ordinary behavior gets read.
  • The explanatory perspectives: behavioral (learned responses), psychodynamic (unconscious conflict), humanistic (blocked growth), cognitive (maladaptive interpretation), evolutionary (once-adaptive responses), sociocultural (conditions people live in).
Study spotlight

Study: Rosenhan (1973)

Eight people with no psychiatric history presented at hospitals reporting one symptom: a voice saying isolated words such as empty and thud. Otherwise their histories were truthful. Every one was admitted. Once inside they said the voice had stopped and behaved as they normally would, taking notes openly. Staff did not identify them; discharge came after an average of about 19 days, with diagnoses recorded as in remission rather than as mistakes.

Reading the research: A covert participant-observation field study, no control condition, no random assignment, outcome measures settled on after the fact. Later archival work, including Cahalan's 2019 investigation, raised serious doubts about how the cases were reported. So the lesson is double: labels are sticky, and a vivid study needs transparency and replication before it becomes a fact.

Apply it

A school team sees a new file: a student diagnosed with obsessive-compulsive disorder. What does the label tell them, and what does it not?

It tells them a pattern met specific criteria, intrusive obsessions and repetitive compulsions that cause distress and eat into the day, and it points to treatments with evidence behind them. It does not predict grades, honesty, or classroom behavior, and it explains nothing they see. Under the diathesis-stress model the diagnosis marks a vulnerability whose expression depends on load, so the team's real lever is the stress side: predictable deadlines, a quiet testing room.

Exam tip: write person-first, "a student diagnosed with OCD," never "an OCD student."

Write it

A classmate says this study proves psychiatric diagnosis is meaningless. Respond in two to four sentences, using the design.

Show a model response

It does not prove that. The pseudopatients deliberately reported a symptom they did not have, so the study shows that clinicians trust what patients tell them, not that the criteria fail. There was no control condition either, and later archival work questioned how the cases were described. What it does show is a labeling effect: once recorded, the diagnosis shaped how ordinary behavior was read.

Why it earns the point: it names a design feature (deception in the input, no control condition) before drawing the narrower conclusion the evidence supports.

Lesson 5.4 · Unit 5 · CED topic 5.4

Anxiety, obsessive-compulsive, and trauma- and stressor-related disorders

Fear is useful equipment. These diagnoses describe it running when nothing is there, running long after the danger has passed, or running so hard that avoiding the fear costs more than the feared thing would. That last part, avoidance, is what turns a bad feeling into a disorder, and it is the hinge the best treatments pull on.

Key terms
  • Generalized anxiety disorder and panic disorder: persistent, hard-to-control worry across many domains for at least six months; versus recurrent unexpected panic attacks plus ongoing worry about the next one.
  • Specific phobia, social anxiety disorder, agoraphobia: marked fear of a particular object or situation; fear of being scrutinized; fear of places that would be hard to leave if panic struck.
  • Obsessive-compulsive disorder: intrusive, unwanted obsessions and repetitive compulsions performed to relieve them. DSM-5-TR groups hoarding disorder and body dysmorphic disorder in the same chapter.
  • Post-traumatic stress disorder: after exposure to trauma: intrusive memories, avoidance, negative shifts in mood and thinking, and hyperarousal, past one month. Acute stress disorder is the same picture inside that first month.
  • Conditioning explanations: classical conditioning attaches fear to a neutral cue; avoidance is then negatively reinforced, since escaping removes distress and the fear is never disconfirmed. Observational learning spreads it secondhand.
  • Cognitive explanations: catastrophizing reads a normal sensation as disaster; attentional bias keeps threat cues in the foreground.
Study spotlight

Study: Foa and colleagues (2005)

One hundred seventy-one women with chronic post-traumatic stress disorder were randomly assigned to prolonged exposure, to prolonged exposure plus cognitive restructuring, or to a waitlist. Prolonged exposure has two parts: revisiting the traumatic memory in a controlled way with a therapist, and gradually approaching safe situations that have been avoided. Assessors who did not know the assignment rated symptom severity. Both active treatments produced large reductions compared with waiting, and adding the cognitive component did not improve on exposure alone.

Reading the research: A randomized controlled trial. The independent variable is treatment condition; the dependent variable is a standardized interview score for PTSD symptom severity. Participants gave informed consent, could withdraw, and the waitlist group was offered treatment afterward. The limitation is the sample, women recruited at specialty clinics, so the result is strongest for the group studied.

Apply it

Two people dislike dogs. Priya crosses the street to avoid one. Tomás has changed his route to work, turned down a job, and stopped visiting his sister.

Priya has ordinary fear. Tomás meets the DSM-5-TR pattern for specific phobia: fear out of proportion to actual danger, active avoidance, and, the deciding criteria, distress and impairment across work and family for months. The diagnosis rests not on the strength of the feeling but on the functional cost. The best-supported treatment is exposure therapy, because it breaks the avoidance loop that keeps the fear from being disconfirmed.

Exam tip: when an item asks whether a behavior is a disorder, look for impairment, not intensity.

Write it

Identify at least one ethical guideline the researchers applied in this study, and explain how the procedure shows it.

Show a model response

They protected participants from harm and honored the right to withdraw. Because exposure means revisiting a traumatic memory, the researchers had participants consent after being told what treatment involved, delivered it with a trained therapist, and let anyone stop without penalty. They also treated the waitlist group afterward rather than leaving people with a serious diagnosis untreated.

Why it earns the point: part D asks for a guideline plus the procedural step that demonstrates it, not just the word "consent."

Lesson 5.5 · Unit 5 · CED topic 5.4

Depressive and bipolar disorders

Sadness is a response to something. The depressive disorders are different: a persistent change in mood, energy, sleep, appetite, and self-evaluation that outlasts its trigger or arrives without one. The bipolar disorders add the other pole: episodes of elevated mood and energy just as far from an ordinary good day.

These are treatable conditions that carry real risk. Warning signs worth acting on include withdrawing from people, giving away belongings, talking about being a burden, and a sudden calm after a long low stretch. If you notice them in yourself or a friend, tell an adult you trust: help is available, and in the US the 988 Suicide and Crisis Lifeline takes calls and texts at any hour.

Key terms
  • Major depressive disorder: two weeks or more of depressed mood or lost interest, plus changes in sleep, appetite, energy, concentration, or self-worth, causing impairment.
  • Persistent depressive disorder: depressed mood more days than not for at least two years in adults; often milder, always longer.
  • Bipolar I and bipolar II disorder: bipolar I requires at least one manic episode: about a week of elevated or irritable mood with markedly increased energy and impaired functioning. Bipolar II pairs shorter, less disruptive hypomanic episodes with major depressive episodes.
  • Seasonal pattern: a specifier for episodes recurring at the same time of year, not a separate diagnosis.
  • Rumination and the negative cognitive triad: turning distress over without resolving it; Beck's negative views of self, world, and future.
  • Explanatory style and diathesis-stress: a pessimistic style blames bad events on internal, stable, global causes. Vulnerability plus stress produces episodes; sleep, exercise, and connection are protective.
Study spotlight

Study: Seligman & Maier (1967); Maier & Seligman (2016)

Dogs were assigned to three conditions: shocks they could switch off, the same shocks with no way to stop them, and no shocks. Later all were placed where escape was simple. Animals that had been able to control the shocks escaped readily; many that had not stayed put and endured it. Seligman called this learned helplessness. In 1978, with Abramson and Teasdale, he recast the mechanism as attributional style. In 2016 Maier and Seligman reversed the original claim on neuroscience evidence: passivity is the default reaction to prolonged aversive events, and what is learned is control.

Reading the research: An animal experiment with prior controllability as the independent variable and escape latency as the dependent variable. The groups were small, and the leap to human depression is an inference, not a finding. A study inflicting inescapable shock would face far stricter animal-care review today.

Apply it

After failing a chemistry test, Nadia says: "I'm just bad at science. I've always been. I'm going to fail everything."

That is a pessimistic explanatory style on all three dimensions. Internal: the cause is her, not the test. Stable: "always been" makes it permanent. Global: "everything" spreads one result across her whole life. A therapist would target the global dimension first: it is easiest to disconfirm with evidence she already has, her other grades. The restructured thought keeps the honesty and drops the overreach: "I did badly here because I skipped the practice problems, and I can change that."

Exam tip: learn the three dimensions as a checklist. Items often give a quotation and ask which dimension it shows.

Write it

Explain how the 2016 revision changes what the original 1967 study can be said to show.

Show a model response

The 1967 result still stands: animals with no control over shocks later failed to escape, while animals that had control escaped. What changed is the interpretation. The original account said helplessness was learned; the 2016 review argues passivity is the unlearned default and the controllable-shock group learned control. The study therefore shows an effect of controllability, not the acquisition of helplessness.

Why it earns the point: it separates the finding from the theory built on it, which is exactly the move an AAQ rewards when a study has been reinterpreted.

Lesson 5.6 · Unit 5 · CED topic 5.4

Schizophrenia spectrum disorders

Schizophrenia is a disorder of how experience is put together: perception, belief, thought, and motivation all shifting at once. It is also the diagnosis most distorted in popular use. It is not dissociative identity disorder, it does not mean a person is dangerous, and it usually arrives gradually, in late adolescence or early adulthood, rather than all at once.

Key terms
  • Positive symptoms: additions to experience: delusions (fixed beliefs held despite contrary evidence), hallucinations (perceptions with no stimulus, most often auditory), disorganized speech and behavior.
  • Negative symptoms: subtractions: flat affect, avolition (little goal-directed activity), alogia (sparse speech), anhedonia. These respond least well to medication.
  • Catatonia and the prodromal phase: marked motor disturbance, from immobility to purposeless movement; and the earlier period of withdrawal and unusual beliefs that precedes a first episode.
  • The dopamine hypothesis: excess dopamine activity in certain pathways. Medications blocking dopamine receptors reduce positive symptoms, which supports it without proving it.
  • Brain and prenatal risk factors: enlarged ventricles and reduced gray matter on average across groups; prenatal viral exposure, malnutrition, and birth complications raise risk modestly.
  • Expressed emotion and diathesis-stress: high criticism, hostility, or over-involvement at home predicts relapse. Genes set vulnerability; environment decides whether it is expressed.
Study spotlight

Study: Gottesman (1991)

Gottesman pooled dozens of European twin, family, and adoption studies into one picture of risk. He compared concordance (the probability that if one relative is diagnosed, the other is too) across degrees of genetic relatedness. For identical twins concordance was roughly 48%; for fraternal twins about 17%; for the general population close to 1%. Children with two parents diagnosed with schizophrenia carried similarly raised risk even when they were raised in other households.

Reading the research: A meta-analysis of behavioral-genetic studies, not an experiment: nobody assigned anyone genes, so these are measured associations. The identical-twin figure carries the argument twice: far above the fraternal rate, so genetic vulnerability is real; far below 100%, so genes alone are not sufficient. That is diathesis-stress in one number. The limitation is that identical twins share more similar environments as well as identical genes.

Apply it

A bar chart shows lifetime risk of a schizophrenia diagnosis by relationship to a diagnosed person. Four bars: general population, just under 1%; sibling, about 9%; fraternal twin, about 17%; identical twin, about 48%. Write the data-analysis answer.

The bars rise with genetic relatedness, and that ordering is the finding: an identical twin's risk is close to three times a fraternal twin's, even though fraternal twins share a womb and a household too. The gap between the two twin bars is the genetic evidence. The gap between the tallest bar and 100% is the environmental evidence: about half the identical twins of a diagnosed person are never diagnosed, so an inherited vulnerability needs something else to become a disorder.

Exam tip: data-analysis items reward reading what the display does not show. A bar chart cannot tell you which environments matter.

Write it

Describe what a concordance rate of about 48% for identical twins indicates about the causes of schizophrenia.

Show a model response

It indicates that when one identical twin is diagnosed, the other is diagnosed a little less than half the time. Because identical twins share essentially all their genes, a rate far above the roughly 1% population risk shows a substantial genetic contribution. Because it is well short of 100%, genes cannot be the whole cause; environmental stressors must also be involved, which is what the diathesis-stress model predicts.

Why it earns the point: part C asks what a statistic indicates, so it reads the number in both directions instead of restating it.

Lesson 5.7 · Unit 5 · CED topic 5.4

Neurodevelopmental, dissociative, somatic symptom, eating, and personality disorders

This lesson collects the remaining DSM-5-TR chapters the exam samples from. They have little in common clinically, so learn each by the feature that distinguishes it from its nearest neighbor: that boundary is exactly what a multiple-choice item tests.

Key terms
  • Neurodevelopmental disorders: attention-deficit/hyperactivity disorder: persistent inattention and/or hyperactivity-impulsivity beginning in childhood, across settings. Autism spectrum disorder: persistent differences in social communication with restricted, repetitive behaviors and interests, evident early.
  • Dissociative disorders: dissociative amnesia is an inability to recall important autobiographical information, beyond ordinary forgetting. Dissociative identity disorder involves two or more distinct personality states with gaps in recall; its prevalence and origins are contested among clinicians.
  • Somatic symptom disorder: distressing physical symptoms plus disproportionate thoughts, feelings, and time spent on them. The symptoms are real; the diagnosis concerns the response.
  • Feeding and eating disorders: anorexia nervosa: restriction leading to significantly low body weight, intense fear of weight gain, disturbed body image. Bulimia nervosa: recurrent binge eating with compensatory behavior. Binge-eating disorder: binges without it.
  • Personality disorder clusters: enduring, inflexible patterns across situations: Cluster A odd or eccentric, Cluster B dramatic or erratic, Cluster C anxious or fearful.
  • Two Cluster B diagnoses: antisocial personality disorder: a long pattern of disregarding others' rights, deceit, impulsivity, little remorse. Borderline personality disorder: instability in relationships, self-image, and emotion, with impulsivity.
Study spotlight

Study: Becker and colleagues (2002)

Adolescent girls in a rural Fijian province completed a standard questionnaire on eating attitudes in 1995, shortly after television reached the region, and a second sample of about the same size completed it in 1998. Scores in the range indicating disordered eating attitudes rose from roughly 13% of the sample to about 29%. Interviews found girls describing a wish to reshape their bodies and naming characters in imported Western programs as their standard, in a culture that had not previously prized thinness.

Reading the research: A naturalistic prospective study built from two cross-sectional samples. Television exposure was measured, not assigned, so this is correlational: the late 1990s also brought tourism, cash work, and schooling changes, any of which could carry the effect. A second limitation is measurement: a questionnaire validated in Western samples may not mean the same thing in Fiji.

Apply it

Sort three vignettes. (1) A man reports severe stomach pain; tests find no cause, and he spends hours daily researching it and visiting clinics. (2) A teenager restricts food, is at a significantly low weight, and describes herself as too large. (3) A woman's relationships swing between idealizing and rejecting the same person, with rapid mood shifts and impulsive decisions.

The first fits somatic symptom disorder: the pain may well be real, and the diagnosis rests on the disproportionate response. The second fits anorexia nervosa: restriction, significantly low weight, and disturbed body image together. The third fits borderline personality disorder, a Cluster B pattern of instability in relationships, self-image, and emotion. Each describes a person with a diagnosis, not a diagnosis with a person attached.

Exam tip: distractors are almost always the neighbouring diagnosis. Name the one criterion that rules the neighbour out.

Write it

Identify the research design used in this study and explain one reason it cannot establish that television caused the change.

Show a model response

This is a correlational, naturalistic study using two cross-sectional samples three years apart. Nobody was randomly assigned to watch television, so any variable that also changed between 1995 and 1998 (tourism, incomes, schooling) could explain the rise in scores. Without random assignment or a comparison region that stayed without television, the design supports an association, not a cause.

Why it earns the point: it names the design in the exam's words and then names a specific confound rather than saying "correlation is not causation."

Lesson 5.8 · Unit 5 · CED topic 5.5

Treatment I: psychodynamic, humanistic, and behavioral therapies

Every therapy is a theory of disorder turned into a procedure. If the trouble is buried conflict, you excavate. If it is a person blocked from growing, you supply the conditions for growth. If it is a learned response, you arrange new learning. Read any technique backwards and you can name the theory that produced it.

Key terms
  • Psychoanalysis and psychodynamic therapy: bringing unconscious conflict into awareness through free association and dream interpretation, so that insight can loosen its hold.
  • Transference and resistance: redirecting feelings about important people onto the therapist; blocking on threatening material. Both are treated as useful signals, not obstacles.
  • Insight versus behavior therapies: change through understanding why, versus change through new learning. Behavior therapy does not assume the symptom has a hidden meaning.
  • Client-centered therapy: Rogers's person-centered approach: active listening, genuineness, empathy, and unconditional positive regard supply the conditions in which a person moves forward on their own.
  • Systematic desensitization and exposure: a fear hierarchy paired with relaxation, which counterconditions the response; or graded contact with the feared stimulus without escape, so avoidance stops being reinforced.
  • Aversive conditioning, token economy, applied behavior analysis: pairing an unwanted behavior with an unpleasant stimulus; earning tokens exchangeable for reinforcers; systematic reinforcement-based programs, widely used and debated over how goals are set.
Study spotlight

Study: Smith & Glass (1977)

Rather than run one more trial, Smith and Glass pooled 375 controlled studies comparing people receiving psychotherapy with untreated controls. Each study was converted to a common measure of how far the treated group's average outcome sat above the control group's, in standard deviations, so different outcome measures could be added together. The average came to about 0.68 standard deviations: the typical treated client finished better off than roughly 75% of untreated controls. Differences between the major therapy types were small next to the difference between therapy and none.

Reading the research: A meta-analysis, one of the first in psychology. The unit of analysis is the study, not the person, and the statistic is an average effect size. Two limitations: publication bias, since null studies sit in file drawers and the pooled estimate is inflated as a result; and uneven quality among the studies pooled.

Apply it

Match each technique to the theory that produced it: (1) a client says whatever comes to mind, uncensored; (2) the therapist reflects feeling back without judging; (3) a client earns chips for self-care, exchangeable for privileges; (4) a client ranks feared situations and works up the list while staying relaxed.

One is free association: psychoanalytic. Two is active listening with unconditional positive regard: humanistic. Three is a token economy: behavioral, built on operant conditioning. Four is systematic desensitization: behavioral, built on classical conditioning. For a specific phobia, choose four: the fear is a conditioned response, and desensitization extinguishes it directly, which is where the outcome evidence is strongest.

Exam tip: if a technique changes what a client does, it is behavioral; if it changes what a client understands, it is an insight therapy.

Write it

Describe what an average effect size of about 0.68 standard deviations indicates about psychotherapy.

Show a model response

It indicates that the average treated client ended up about two-thirds of a standard deviation better off than the average untreated control: enough to place the typical treated person above roughly 75% of the untreated group. That is a substantial, not a trivial, difference. It does not say which therapy is best, because it averages across many types, and the pooled estimate may be inflated by publication bias.

Why it earns the point: it translates the statistic into a comparison between groups and then names what the number cannot settle.

Lesson 5.9 · Unit 5 · CED topic 5.5

Treatment II: cognitive, cognitive-behavioral, and group approaches

The cognitive therapies start from a claim you can test on yourself: the event is not what upsets you, your reading of it is. Two students get the same B−. One files it and moves on; the other concludes she is not cut out for this. Same event, different belief, and the belief is the part that can be examined and changed.

Key terms
  • Cognitive therapy and restructuring: catch the automatic thought, test it against evidence, replace it with something more accurate. In depression the target is Beck's triad: negative views of self, world, and future.
  • Rational emotive behavior therapy and the ABC model: Ellis's sequence of activating event, belief, consequence. The therapist disputes the belief, since the event is already over.
  • Cognitive-behavioral therapy: thought work plus behavioral assignments such as exposure or activity scheduling. The most extensively tested approach across disorders.
  • Dialectical behavior therapy: a CBT adaptation holding acceptance and change together, teaching distress tolerance, mindfulness, and emotion regulation. Developed for people diagnosed with borderline personality disorder.
  • Group, family, and couples therapy: the relationships are the medium, not just the setting. Self-help groups add peer support and modeling at low cost, without a clinician.
  • Alliance, evidence-based practice, cultural competence: the collaborative bond between client and therapist predicts outcome across approaches; evidence-based practice combines research, clinical expertise, and the client's values and culture.
Study spotlight

Study: Rush, Beck, Kovacs & Hollon (1977)

Forty-one outpatients with moderate to severe depression were randomly assigned to twelve weeks of cognitive therapy or to the antidepressant imipramine. Severity was rated before, during, and after treatment on standardized depression inventories, one completed by the patient and one by a clinician. The cognitive therapy group showed greater reduction in depression scores, and fewer of them dropped out, among the first demonstrations that a talking therapy could match or beat medication on a standardized measure.

Reading the research: A randomized clinical trial. The independent variable is treatment type; the dependent variable is score on standardized depression inventories. Patients consented and were clinically monitored throughout. The limitations are real: a small sample at a single site, no placebo arm, no blinding of patients to which treatment they were getting, and drug dosing that later trials questioned.

Apply it

Jamal is not invited to a party and thinks: "Nobody likes me. I'm always going to be the one left out." Work it through the ABC model.

A, the activating event, is not being invited to one party. B, the belief, is the distortion: overgeneralization, turning one omission into "nobody" and "always." C, the consequence, is low mood plus a decision to stop texting people, which will manufacture more evidence for the belief. The restructured thought stays honest: "I wasn't invited to this one, and that stings. Two friends texted me this week." A therapist disputes B, since A cannot be changed and C follows from B.

Exam tip: application items about cognitive therapy almost always hinge on locating B. The upsetting event is never the answer.

Write it

Identify the research method used in this study and explain one design feature that limits what can be concluded from it.

Show a model response

The researchers used an experiment: a randomized clinical trial, since patients were randomly assigned to cognitive therapy or to medication. One limiting feature is the absence of a placebo condition: with no group receiving an inactive pill and the same attention, the difference between treatments cannot be separated from expectancy. The small single-site sample also makes the estimate imprecise.

Why it earns the point: it names the method, then names a specific missing control rather than saying the study was "not accurate."

Lesson 5.10 · Unit 5 · CED topic 5.5

Biomedical treatment and the ethics of care

Biomedical treatments change the nervous system directly. That makes them powerful and makes the consent conversation unavoidable: side effects are real, the person taking the medication lives with them, and a competent adult may refuse. The history of this field is why those protections are written down.

Key terms
  • Antipsychotic medication: reduces positive symptoms largely by blocking dopamine receptors. Long-term use risks tardive dyskinesia, involuntary movements that can persist after the drug is stopped.
  • Antidepressants and anxiolytics: selective serotonin reuptake inhibitors block reuptake, leaving more serotonin at the synapse; benefits build over weeks. Anxiolytics dampen arousal quickly and carry dependence risk.
  • Lithium and mood stabilizers: reduce the frequency and severity of episodes in bipolar disorder, with blood levels monitored because the effective and unsafe ranges sit close together.
  • ECT and rTMS: electroconvulsive therapy, given under anesthesia, is reserved for severe depression that has not responded to other treatments; memory effects are the main cost. Repetitive transcranial magnetic stimulation is a non-invasive alternative. The psychosurgery era is discarded history.
  • Deinstitutionalization: moving care out of large state hospitals. Without funded community services, many people were left without care rather than better care.
  • Placebo effect, double-blind, consent: improvement produced by expectation alone; a design in which neither participant nor assessor knows the assignment; and the treatment rights that follow, including informed consent and the right to refuse.
Study spotlight

Study: Trivedi and colleagues; Rush and colleagues (2006)

The STAR*D trial enrolled about 2,900 outpatients with major depressive disorder at clinics chosen to resemble ordinary practice rather than a research ward. Everyone began on the same medication, citalopram, given openly rather than blind. Roughly a third reached remission: scores in the non-depressed range on a standardized rating scale, not merely improvement. Those who did not moved to a second step, switching medications or adding one, then to third and fourth steps. Each successive step produced smaller additional gains, with cumulative remission approaching two-thirds.

Reading the research: A large multi-site sequenced treatment trial. The dependent variable is remission on a standardized rating scale, and the sequence is the design's whole point, since real patients who do not respond get a second option. The limitation is control: the first step was open-label with no placebo arm, so how much of that third would have remitted on an inactive pill cannot be separated out.

Apply it

A family reads that a new supplement "beat depression in 8 weeks" in a company study where everyone took it. Explain why that is not enough.

With one group and no comparison, improvement has three explanations besides the supplement: expectancy, ordinary remission over eight weeks, and regression toward the mean, since people enroll when they feel worst. A double-blind placebo-controlled design fixes this: an inactive pill supplies the baseline, and keeping both participant and rater unaware of the assignment stops expectancy from moving the scores. Two safeguards such a trial must carry: consent that discloses the chance of receiving placebo, and monitoring that moves anyone who worsens into active treatment.

Exam tip: scientific investigation items love this stem. The missing element is usually the control group.

Write it

Describe what a remission rate of roughly one third at the first treatment step indicates about treating major depressive disorder.

Show a model response

It indicates that about one patient in three reached the non-depressed range on the first medication tried, so most did not: a first prescription is a starting point, not a solution. Because later steps helped progressively fewer people, treatment-resistant depression is common. And since the first step had no placebo group, part of that third may reflect expectancy or natural recovery.

Why it earns the point: part C wants what the statistic indicates, so it reads the number, the trend across steps, and the limit the missing control imposes.

Lesson 5.11 · Unit 5 · Research methods

Reading a research article: putting the AAQ skills together

The Article Analysis Question hands you one research summary and twenty-five minutes. Everything it asks for is already in the text, in a predictable place. So the skill is not reading faster; it is knowing which paragraph holds which answer.

Key terms
  • The four sections: the abstract summarizes; the method gives design, participants, and procedure; the results carry the statistics; the discussion interprets, and is where authors overreach.
  • Identifying the design: a manipulated variable plus random assignment means experiment; measured variables only means correlational; one person means case study; pooled studies means meta-analysis.
  • Locating operational definitions: the sentence saying how a construct was measured, usually under measures. It names an instrument, a count, or a scale.
  • Reading a reported statistic: translate it into a statement about groups: a mean difference, a percentage, a correlation such as \(r = -0.32\), an effect size such as d = 0.8.
  • Finding the ethics statement: a sentence near the participants paragraph about consent, debriefing, withdrawal, confidentiality, or board review.
  • Generalizability and preregistration: the participants paragraph tells you how many, who, and from where. Preregistration posts the hypothesis and analysis plan before data collection, so the analysis cannot be chosen to fit the result.
Study spotlight

Study: Carney, Cuddy & Yap (2010); Ranehill and colleagues (2015)

Forty-two participants were randomly assigned to hold either expansive, open postures or contracted, closed ones for two minutes. The high-power group showed hormonal changes in saliva samples, took the riskier option more often on a gambling task, and reported feeling more powerful. Five years later a preregistered replication ran the same manipulation with 200 participants and a tighter procedure. It reproduced the self-reported feeling of power, but found no hormonal effect and no effect on risk-taking.

Reading the research: Two experiments with the same independent variable and different dependent variables surviving. Random assignment in both makes them causal designs, so the disagreement is not about design type but about precision. A 42-person study gives a noisy estimate; 200 participants and a preregistered analysis plan give a tighter one. A single striking result is a hypothesis, not a fact.

Apply it

Map the six AAQ parts onto the sections of the replication summary above.

A, the method, comes from the design sentence: participants randomly assigned to a posture, so it is an experiment. B, an operational definition, comes from the measures: feeling powerful was a self-report rating. C comes from the results, no reliable hormone difference between conditions. D, an ethical guideline, comes from the procedure: consent before an unusual physical task. E, generalizability, comes from the participants paragraph. F comes from you: the results refute the hormonal claim while supporting the felt-power one.

Exam tip: answer in order and label each part. Readers score A through F as separate rows; an unlabeled paragraph loses points it earned.

Write it

Answer parts A, E, and F for the 2015 replication: identify the method, judge generalizability with participant evidence, and explain how the results support or refute the claim that posture changes how powerful people feel and act.

Show a model response

A. The researchers used an experiment, since participants were randomly assigned to a high-power or low-power posture. E. Generalization is moderate: 200 adult volunteers in a single laboratory is large enough for a precise estimate but still one setting, so the result should not be stretched to negotiations or interviews outside the lab. F. The results partly refute the original claim. They support a psychological effect, expansive posers reported feeling more powerful, but refute the physiological and behavioral claims, since neither the hormone measures nor the risk task differed by condition.

Why it earns the point: each part is labeled and answered on its own terms, and F splits the claim instead of accepting or rejecting all of it.

Unit 5 review · 10 multiple-choice

Unit 5 review: Mental and Physical Health

Ten questions spread across all eleven lessons in the exam's three styles (concept application, data analysis, and scientific investigation) with DSM-5-TR criteria, treatment evidence, and the design questions that decide what a clinical result is worth.

Multiple choice

  1. For three weeks Priya works long hours on a demanding project, feels wired but functional, and sleeps badly. The morning after she submits it she comes down with a heavy cold. Which stage of the general adaptation syndrome does the cold best illustrate?

    Alarm is the opening burst (sympathetic activation, adrenaline, a body mobilized in seconds), and it belongs at the start of the three weeks, not at the end. By the time the cold arrives the stressor is over, which is the giveaway that a later stage is being tested.

    Correct. Selye's third stage is where the bill comes due: cortisol held high through the resistance phase suppresses immune function, so illness tends to surface once the demand lifts. Note the coping split the lesson uses: a fixed revision schedule is problem-focused because it changes the stressor, while a run is emotion-focused because it changes her reaction.

    Resistance is the plateau during the three weeks: costly but functional, exactly the "wired but functional" stretch the scenario describes. The question asks about the cold, which comes after that plateau ends, so resistance names the wrong part of the sequence.

    Eustress is arousal that sharpens performance rather than swamping it, and it may well describe her first week. It is a label for a quality of stress, not a stage of the general adaptation syndrome, so it cannot answer a question about which stage the illness shows.

  2. Described study

    Researchers recruited 240 adults through a well-being website and randomly assigned each one to one of two one-week writing exercises: listing three good things each evening along with what caused them, or writing each evening about early childhood memories. Both groups were asked to write for the same number of minutes and were told the exercise was expected to help. Happiness and depressive-symptom questionnaires were completed before the week, immediately after, and again at one, three, and six months.

    What is the main purpose of the early-memories writing condition?

    That question could be asked of the data, but it is not why the condition is there. A control condition is defined by its role in the comparison, it supplies the baseline the treatment is measured against, not by an interest in the control activity itself.

    Power comes from how many participants are in the study, and splitting them across conditions is how an experiment is built, not a way of manufacturing power. If sample size were the aim, the researchers would simply put all 240 in one group and have no comparison at all.

    An active comparison condition drawn from a rival theory is a legitimate design, but nothing here presents early-memory writing as a treatment anyone predicted would work. It was chosen because it looks and feels like the exercise while lacking the ingredient under test.

    Correct. A control that does nothing leaves expectancy and time-on-task free to explain any improvement; this one matches both and withholds only the active ingredient, crediting causes and noticing good things. That is what makes the comparison clean, and it is the design point a scientific investigation item usually targets.

  3. A ninth grader checks his locker four or five times before every class. He knows it is locked, the checking makes him late, and he has started leaving home early to make room for it. He describes feeling ashamed and unable to stop. A counselor is weighing whether this pattern crosses the threshold for a diagnosis. Which criteria are doing most of the work?

    Correct. Of the four criteria clinicians weigh (distress, dysfunction, deviation, danger) two are clearly present: he suffers, and the behavior costs him time and punctuality. That functional cost, not the strangeness of the behavior, is what moves a habit toward a diagnosis such as obsessive-compulsive disorder, and it is why a person can be unusual without being unwell.

    Deviation from a norm is the weakest of the four criteria on its own: plenty of statistically unusual behavior is harmless, and what counts as deviant varies by culture and setting. Using it alone is how ordinary difference gets pathologized, which is the error the criteria are designed to prevent.

    Danger means risk of physical harm to the person or to others, and nothing in the scenario describes any. It is also the rarest of the four criteria in practice: most people who meet criteria for a psychological disorder pose no danger to anyone, and assuming otherwise is the stereotype that drives stigma.

    Comorbidity describes two or more diagnoses co-occurring, such as a depressive disorder alongside an anxiety disorder. It is a real and common feature of clinical practice, but it is not a criterion for deciding whether a single pattern is a disorder in the first place.

  4. Tomás is afraid of dogs. Each time he sees one he crosses the street, and his anxiety drops within seconds. He has now rerouted his commute and stopped visiting a friend who has a dog. Which principle best explains why the avoidance persists and why the fear does not fade?

    Relief is the removal of something unpleasant, not the addition of something pleasant, so this is negative rather than positive reinforcement. The two are the most confused pair in the unit: both raise behavior, and the sign refers to whether a stimulus is added or taken away.

    Punishment decreases the behavior it follows, and Tomás's avoidance is increasing, so the term has the wrong direction built in. Notice that the scenario reports a behavior becoming more frequent and more elaborate, which rules out punishment before any other reasoning starts.

    Correct. The avoidance is maintained because it works in the short run, and that short-run success is the trap: each crossing removes anxiety immediately, while guaranteeing he never learns that a leashed dog on the sidewalk passes without incident. Breaking that loop is exactly what exposure therapy is for, which is why it is the best-supported treatment for specific phobia.

    Classical conditioning plausibly explains how the fear was acquired, but the roles are assigned wrongly: a frightening event would be the unconditioned stimulus and the dog the conditioned stimulus. And acquisition is not the question asked: the item asks what keeps the fear alive now, which is the operant half of the two-factor account.

  5. After a poor audition, Wren says: "I choked because I'm just not a performer. I never have been and I never will be. And it's the same with everything I try." Which part of her statement expresses the global dimension of a pessimistic explanatory style?

    This is the event being explained, not the explanation of it. Attributional style is about the causes a person assigns, so the phrase that describes the outcome cannot be the one that carries a dimension.

    This is the internal dimension: the cause is located in her rather than in the situation, the piece, or the panel. Internal, stable, and global are three separate dimensions, and items like this one ask you to keep them apart.

    This is the stable dimension: the cause is treated as permanent rather than temporary. It is the one most often mistaken for global, but stability is about time, while global is about breadth across situations.

    Correct. Global means the cause is taken to apply everywhere, so one bad audition becomes evidence about her whole life. A cognitive therapist would often dispute the global claim first, because it is the easiest to test against evidence Wren already has: the other things she does competently.

  6. Described graph

    A bar chart shows the lifetime risk of a diagnosis of schizophrenia by genetic relationship to a person diagnosed with it, pooled across dozens of European twin, family, and adoption studies (Gottesman, 1991). Risk in percent runs up the vertical axis. Four bars are shown: general population, just under 1%; sibling, about 9%; fraternal (dizygotic) twin, about 17%; identical (monozygotic) twin, about 48%.

    Which conclusion is best supported by the display?

    A single-gene disorder would produce concordance near 100% in identical twins, who share essentially all their genes. The chart shows about 48%, which is the strongest single piece of evidence against a one-gene account and for many genes of small effect plus environmental contributions.

    Correct. Read the display in both directions, which is what a data analysis item rewards: 48% against 17% for pairs who share a womb and a household is the genetic evidence, because identical twins differ from fraternal twins mainly in genes; 48% against 100% is the environmental evidence. Together they are the diathesis-stress model in one chart.

    The bars are not close, 17% is nearly double 9%, and even if they were, a chart of relatedness categories cannot isolate shared environment. That comparison needs designs the display does not contain, such as identical twins reared apart or adoption studies.

    The number is read correctly and then explained backwards. Identical and fraternal twin pairs are both typically raised in the same household, so a shared household cannot explain why one group's risk is nearly three times the other's; what differs between them is genetic similarity.

  7. Described data

    Adolescent girls in a rural Fijian province completed a standard questionnaire on eating attitudes in 1995, shortly after television became available in the region, and a separate sample of about the same size completed it in 1998 (Becker and colleagues, 2002). The table gives the percentage of each sample scoring in the range that indicates disordered eating attitudes.

    Survey yearPercent scoring in the at-risk range
    1995, television newly availableabout 13%
    1998about 29%

    Which statement is best supported by these data?

    Correct. Both halves are supported: 29% against 13% is more than double, and no one was randomly assigned to watch television, so the design is correlational. The honest reading names a plausible alternative: the late 1990s also brought tourism, cash work, and schooling changes to the province, any of which could carry the effect.

    Two errors in one sentence. The causal verb is not licensed by a study with no random assignment, and the arithmetic is misreported: 13% to 29% is a rise of 16 percentage points, which is a relative increase of more than 120%, not "a 16% increase."

    A screening questionnaire measures attitudes, not diagnoses. Scoring in an at-risk range means a clinician might want to look further; a diagnosis of anorexia nervosa requires restriction leading to significantly low body weight, intense fear of weight gain, and disturbed body image, assessed by a clinician.

    The stimulus says a separate sample of about the same size completed the 1998 survey, so these are two cross-sectional samples three years apart, not the same girls followed over time. That distinction matters: a repeated cross-section describes a population shift and cannot tell you that any individual's attitudes changed.

  8. Described data

    Adults who met criteria for a specific phobia of dogs were randomly assigned to eight sessions of systematic desensitization, eight sessions of cognitive restructuring alone, or a waitlist. Avoidance was measured with a behavioral approach test scored from 0 (will not enter the room) to 40 (will hold the leash), administered before treatment and again after eight weeks by an assessor who did not know the assignment. Higher scores mean less avoidance.

    GroupnMean beforeMean after 8 weeks
    Systematic desensitization3211.430.6
    Cognitive restructuring only3111.822.1
    Waitlist3011.613.2

    Which statement is best supported by the table?

    Compare each group with the control, not with the best group. Cognitive restructuring gained 10.3 points against the waitlist's 1.6, which is a substantial benefit; it is simply a smaller one than desensitization produced.

    A 1.6-point move on a 40-point scale over eight weeks is close to flat, and it is the reason the waitlist is there, to show what happens without treatment. Reading it as spontaneous recovery inverts the evidence, and it would also have to explain why the treated groups moved six to twelve times as far.

    Correct. Subtract before from after in each row: 30.6 − 11.4 = 19.2, 22.1 − 11.8 = 10.3, and 13.2 − 11.6 = 1.6. The comparison is fair because random assignment left the groups nearly equal at baseline and the assessor was blind to condition, which is why the gap between the two active treatments can be read as the contribution of exposure itself.

    Similar baselines show that random assignment worked, not that the measure is valid. Validity would require the approach-test score to line up with something independent (clinician ratings, or real avoidance outside the clinic), and a table of group means cannot supply that.

  9. Described study

    A company gave a new herbal supplement to 300 adults who reported persistent low mood. Everyone received the supplement, everyone knew what they were taking, and each person completed a standardized depression rating scale at enrollment and again after eight weeks. The group's mean score fell from 18.6 to 13.9. The company's press release states that the supplement "relieves depression in eight weeks."

    Which change to the design would most directly address the strongest objection to that conclusion?

    More participants would pin down the 4.7-point drop more tightly, but the problem is not how precisely the drop is measured: it is that a one-group design offers no way to tell what produced it. Thirty times the sample with no comparison condition buys a precise estimate of an uninterpretable number.

    Correct. With a single group, at least three explanations compete with the supplement: expectancy, ordinary remission over eight weeks, and regression toward the mean, since people enroll when they feel worst. A placebo arm supplies the baseline and double-blinding keeps expectancy from moving either the participants' answers or the raters' scores, and such a trial must also disclose the chance of receiving placebo in the consent form and move anyone who worsens into active treatment.

    Weekly measurement produces a nicer curve and might reveal when change happens, but every point on that curve still belongs to one group with nothing to compare it against. More frequent measurement of the wrong design does not fix the design.

    A clinician interview is generally a stronger outcome measure than self-report, and in an unblinded study it can be worse, since a clinician who knows everyone took the supplement brings expectations too. Either way, improving the dependent variable leaves the missing control group untouched.

  10. A research summary contains this sentence: "Symptom severity was recorded as each participant's total score on a 17-item clinician-administered interview, completed before treatment and again after twelve weeks." Which part of an Article Analysis Question does this sentence most directly answer?

    Correct. An operational definition says how a construct was actually measured, here a total score on a named 17-item interview at two fixed time points, which is what part B is scored on. Writing "symptom severity means how bad the symptoms were" is a dictionary definition and earns nothing; the instrument and the score are the point.

    Part A wants the design named: experiment, quasi-experiment, correlational study, case study, naturalistic observation, survey, longitudinal design, or meta-analysis. A measures sentence describes an instrument, not a design, and you would look instead for the sentence about assignment to conditions.

    Part C asks you to say what a reported number indicates about the findings: what a mean difference, a percentage, a correlation, or an effect size tells us here. The sentence quoted reports no number at all; it describes how numbers were generated.

    Part E is answered from the participants paragraph: how many people, their ages, sex, culture, how they were recruited, and the setting. The sentence quoted names who administered the interview, which is procedure, not a participant characteristic.

AAQ 1 · 25 minutes · 7 points

Using the source, respond to parts A through F: do the reported results support or refute the claim that memory is reconstructive?

Directions

Read the source, then respond to all six parts in complete sentences, in order, labelling each part with its letter. You may write the parts as separate short paragraphs. Do not bring in a different study; every claim you make about the research must come from this source. Parts A through E are worth one point each and part F is worth two, for seven points in twenty-five minutes.

  • A. Identify the research method used in the study.
  • B. State the operational definition of the dependent variable in the first experiment, the speed estimate.
  • C. Describe what the difference between the mean speed estimate of 40.8 miles per hour and the mean speed estimate of 34.0 miles per hour indicates about the study's findings.
  • D. Identify at least one ethical guideline the researchers applied.
  • E. Explain the extent to which the findings can be generalized beyond the people who took part, using specific evidence about the participants.
  • F. Explain how the study's specific results support or refute the claim that memory is reconstructive. (2 points)
Source

Source: research summary written for this course, based on Loftus and Palmer (1974).

The researchers asked whether the wording of a question put to a witness after an event changes what that witness later reports about the event. Two experiments were run with undergraduate students at an American university, tested in groups in a campus laboratory room.

In the first experiment, 45 students watched seven short films of staged traffic accidents, each running between about five and thirty seconds, drawn from driver-education safety materials. After each film the students wrote an account of the accident in their own words and then answered a set of specific questions about it. One question was critical. Students were asked how fast the cars were going when the collision occurred, and the verb used in that question differed by group: nine students were asked with the verb smashed, nine with collided, nine with bumped, nine with hit, and nine with contacted. Each student answered by writing a speed in miles per hour. The mean estimate was 40.8 miles per hour among students asked with smashed and 34.0 miles per hour among students asked with hit, with the other three verbs producing means that fell between those two figures.

A second experiment asked whether the wording changes the remembered event itself rather than only the answer given at the moment. A further 150 students watched a one-minute film that contained a brief multiple-car accident. Fifty were asked the speed question with the verb smashed, fifty with the verb hit, and fifty were asked no question about speed at all. One week later all 150 returned and answered a new set of questions without seeing the film again. Among these was a question about whether they had seen any broken glass. There was no broken glass anywhere in the film. About 32 percent of the students questioned with smashed reported that they had seen broken glass, compared with about 14 percent of those questioned with hit and about 12 percent of those who had been asked nothing about speed.

Students volunteered for the sessions after being told they would watch films of traffic accidents and answer questions about what they had seen, and they could stop at any point. The films were standard driver-education footage rather than graphic material. Responses were recorded by group rather than by name, and the purpose of varying the verb was explained to the students once their questionnaires had been collected.

Your response
What a reader looks for on this prompt
  • A. The method, named: the word experiment (or laboratory experiment) has to appear. The source gives you the two features that make it one: groups received different verbs while the films were held constant. Describing that procedure without naming the method earns nothing, and "survey" or "correlational study", tempting because the data came from a questionnaire, contradicts the source.
  • B. How the variable was measured here: the speed estimate in miles per hour that each student wrote on the questionnaire after viewing the film. Saying "how fast the students thought the cars were going" is a definition of the construct, not an operational definition; the point comes from naming the written number and its unit.
  • C. Interpret the gap, do not restate it: 40.8 and 34.0 are means for groups who saw identical films, so the 6.8 mile-per-hour difference cannot have come from what was on screen. Say what it indicates: the verb in the question shifted reported speed by almost seven miles per hour. Writing "a mean is an average" defines a statistic instead of interpreting this one.
  • D. A guideline they followed: informed consent (students were told what they would watch and could stop), debriefing (the verb manipulation was explained after the questionnaires were collected), the right to withdraw, or confidentiality of responses. A criticism of the study, that the wording misled people, is an ethical problem, not a guideline applied, and earns nothing.
  • E. Participant evidence first, then the judgment: 45 undergraduates at one American university in the first experiment and 150 in the second, all students, all watching short films in a laboratory rather than witnessing a crash. Cite at least one of those details and then say how far the findings travel. A bare "this cannot be generalized" earns nothing.
  • F. Results plus the concept, both (2 points): cite figures from the source (40.8 against 34.0; 32 percent against 14 and 12 percent) and explain that a stored recording could not be altered by a question asked afterward, so retrieval must rebuild the event from the trace plus later information. One point if you cite the numbers without linking them to reconstruction, or explain reconstruction without the numbers.
Show a 7/7 response

A The researchers used an experiment. Students were assigned to groups that received different verbs in the critical question while the films were held constant, so a variable was manipulated across conditions: a laboratory experiment, not a survey or a correlational study.

B In the first experiment the dependent variable was the speed estimate, operationally defined as the number of miles per hour each student wrote on the questionnaire in answer to the question about how fast the cars were going when the collision occurred. In the second experiment it was a yes-or-no answer to the question about broken glass one week later.

C Both numbers are group means for students who watched exactly the same films, so the 6.8 mile-per-hour difference between 40.8 and 34.0 cannot come from anything they saw. It indicates that a single word in the question changed the reported speed: students asked with the stronger verb reported a crash almost seven miles per hour faster than students asked with the milder one. The mean difference measures the size of that wording effect.

D The researchers applied informed consent. The source reports that students volunteered after being told they would watch films of traffic accidents and answer questions about them, and could stop at any point. They also debriefed participants, explaining the purpose of varying the verb once the questionnaires had been collected.

E Generalization is limited, and the participant details show why. The first experiment used 45 undergraduates at a single American university and the second 150 more from the same population, so the sample was narrow in age, education and country. They also watched short driver-education films in a laboratory, where nothing was at stake, rather than witnessing a real crash. The wording effect has been found with other groups, so the direction probably extends beyond students, but the size of these means and percentages should not be carried over to witnesses of real events.

F The results support the claim that memory is reconstructive. If remembering worked like replaying a recording, the question asked afterward could not touch the stored copy, and every group watching the same film would report the same speed. Instead the mean estimate was 40.8 miles per hour with one verb and 34.0 with another, and a week later about 32 percent of the students asked with the stronger verb reported broken glass that was never in the film, against about 14 percent asked with the milder verb and about 12 percent asked nothing about speed. The broken-glass result is the stronger evidence, because a detail that did not exist was reported as something seen, seven days later, by people whose only difference was one word in an earlier question. That is what reconstruction means: retrieval rebuilds the event from the original trace together with information that arrived afterward, and the post-event wording became part of what these students reported remembering. The misinformation effect names this pattern, and the percentages rise with how much misleading information each group received.

Where the points are earned
  • A: Research method (1): names the method, "an experiment", and justifies the label with the manipulation of the verb against constant films. Naming it is what earns the point; the justification protects against a reader who suspects a lucky guess.
  • B: Operational definition (1): gives the measure as the researchers took it, a number of miles per hour written on a questionnaire in answer to a named question, rather than defining "speed estimate" in the abstract. The second sentence covers the follow-up measure in case a reader treats the broken-glass report as the identified variable.
  • C: Interpreting the statistic (1): interprets the mean difference in context instead of defining a mean: identical films, a 6.8 mile-per-hour gap, therefore an effect of the question's wording on what was reported.
  • D: Ethical guideline (1): identifies informed consent and grounds it in what the source actually reports, students were told what they would watch and could stop, then adds debriefing as a second guideline. Either alone would earn the point.
  • E: Generalizability (1): states a limit and supports it with participant evidence from the source: 45 undergraduates, then 150 more, one American university, laboratory films rather than a witnessed crash. It also separates the direction of the effect from its magnitude, which is the honest version of the judgment.
  • F: Argumentation, evidence (1 of 2): cites specific results rather than gesturing at them: 40.8 against 34.0 miles per hour, and 32 percent against 14 and 12 percent reporting broken glass a week later.
  • F: Argumentation, explanation (2 of 2): links those results to reconstructive memory and uses the concept correctly: a recording could not be altered by a later question, so retrieval must combine the original trace with post-event information, which is the misinformation effect. The response says which result is the stronger evidence and why, which is what separates a full-credit argument from a summary.

AAQ 2 · 25 minutes · 7 points

Using the source, respond to parts A through F: do the reported results support or refute observational learning as an explanation of aggression?

Directions

Read the source, then respond to all six parts in complete sentences, in order, labelling each part with its letter. You may write the parts as separate short paragraphs. Do not bring in a different study; every claim you make about the research must come from this source. Parts A through E are worth one point each and part F is worth two, for seven points in twenty-five minutes.

  • A. Identify the research method used in the study.
  • B. State the operational definition of the dependent variable, imitative physical aggression.
  • C. Describe what the difference between a mean of about 25.8 imitative physically aggressive acts among boys who watched an aggressive male model and a mean of about two among children who watched no model indicates about the study's findings.
  • D. Identify at least one ethical guideline the researchers applied.
  • E. Explain the extent to which the findings can be generalized beyond the children who took part, using specific evidence about the participants.
  • F. Explain how the study's specific results support or refute observational learning as an explanation of aggression. (2 points)
Source

Source: research summary written for this course, based on Bandura, Ross and Ross (1961).

The researchers tested whether young children will reproduce specific aggressive acts performed by an adult they have watched, in the adult's absence and without reward.

The participants were 72 children at a university nursery school in the United States, 36 boys and 36 girls, aged about three to six years (mean age near four years and four months). Each child was first rated for everyday aggressiveness by an experimenter and a teacher who knew the child. Children with similar ratings were matched in groups of three, and one child from each group was placed in each of three conditions: an aggressive model, a non-aggressive model, or no model. Within the model conditions half the children saw an adult man and half an adult woman.

Children in the model conditions were brought one at a time to a room with craft materials at one table and, in a corner, a mallet, toys, and a five-foot inflatable doll weighted to right itself when struck. For about ten minutes the non-aggressive model assembled toys quietly and ignored the doll. Over the same period the aggressive model performed a fixed sequence of distinctive acts against it: laying it on its side and punching its nose, striking its head with the mallet, tossing it in the air, and kicking it about the room, while repeating a short set of aggressive phrases. Every child was then taken to a room of attractive toys and, after a brief play period, told they were being saved for other children: a mild frustration procedure used in all three conditions. Each child was then observed for twenty minutes in a test room containing the mallet, the doll and other toys, with an adult present. Observers behind a one-way mirror recorded behaviour in five-second intervals, giving 240 scored intervals per child, and counted those containing an act reproducing what the model had done.

Children who had watched the aggressive model reproduced those acts far more often than the other two conditions, whose imitative counts were close to zero. Boys who had watched an aggressive male model averaged about 25.8 imitative physically aggressive acts; girls who had watched an aggressive female model averaged about 5.5; children who saw no model averaged about two or fewer. Counting all aggressive behaviour rather than imitation alone, the aggressive-model group produced roughly twice as much as the other groups, and boys imitated physical aggression more than girls did.

The children took part with the permission of their parents. The frustration procedure was kept brief and mild, an adult stayed in the room throughout the test, and the records were coded interval counts rather than identifying accounts.

Your response
What a reader looks for on this prompt
  • A. The method, named: experiment (a laboratory experiment with matched participants). The source hands you the evidence: children were matched on prior aggressiveness and then assigned to three conditions that differed only in what the adult did. "Naturalistic observation" is the trap here: observation through a one-way mirror is how the data were collected, not the design, because the researchers arranged what each child saw.
  • B. How the variable was counted here: the number of five-second intervals, out of the 240 in the twenty-minute session, in which observers recorded the child performing one of the acts the model had performed. Calling it "how aggressive the child was" restates the construct; the point comes from the interval count and the fact that only reproductions of the model's specific acts were scored.
  • C. Interpret the gap, do not restate it: about 25.8 against about two is a difference of roughly an order of magnitude between groups that had been matched on aggressiveness beforehand, so the gap indicates that what the children watched, not who they already were, produced the behavior. Saying "25.8 is bigger than 2" is not an interpretation.
  • D. A guideline they followed: parental consent, protection from harm (the frustration was deliberately brief and mild, an adult stayed in the room), or confidentiality (coded interval counts rather than identifying records). Writing that frustrating children was unethical names a problem rather than a guideline applied, and earns nothing.
  • E. Participant evidence first, then the judgment: 72 children, aged about three to six, all enrolled at one university nursery school in the United States in the early 1960s. Cite at least one of those facts, then say how far the findings travel and note that the target was an inflatable toy built to be struck.
  • F. Results plus the concept, both (2 points): use the numbers (about 25.8 against about two, imitative counts near zero without an aggressive model) and explain that the children performed distinctive acts they had never been taught and were never reinforced for, in the model's absence: acquisition by observation. A full-credit answer also marks the limit: imitation directed at a doll. One point for numbers without the concept, or the concept without numbers.
Show a 7/7 response

A The researchers used an experiment. Children were matched on prior aggressiveness ratings and then assigned to one of three conditions (aggressive model, non-aggressive model, or no model), so the adult's behaviour was manipulated while everything else was held constant.

B Imitative physical aggression was operationally defined as the number of five-second intervals, of the 240 in the twenty-minute test, in which observers behind a one-way mirror recorded the child performing one of the specific acts the model had performed, such as punching the doll's nose, striking it with the mallet, or tossing it in the air.

C Boys who watched an aggressive male model averaged about 25.8 of these acts while children who saw no model averaged about two, so the modelled group produced them more than ten times as often. Because the children had been matched on how aggressive they already were, that difference indicates the effect of what the child watched rather than a difference in the children themselves.

D The researchers applied consent in the form appropriate to children: the source reports that they took part with the permission of their parents. They also protected them from harm, since the frustration procedure was kept brief and mild and an adult stayed in the room throughout.

E Generalization is limited. All 72 participants were three- to six-year-olds at a single university nursery school in the United States, a narrow slice by age, culture and background, and the behavior was measured minutes later in a laboratory room. The target was an inflatable doll built to be struck, not a person, so the results speak to whether young children copy what they see, not to whether watching aggression harms others. That children imitate adult models has held up widely, but these counts are not a measure of real-world aggression.

F The results support observational learning as an explanation of how aggressive behavior is acquired. The children who saw the aggressive model reproduced a fixed sequence of unusual acts (punching the doll's nose, hitting it with a mallet, throwing it in the air) that no one had taught them and for which they received no reward, and they did it with the model gone. That is the core claim of observational learning: behavior can be acquired by watching a model, with no reinforcement of the learner. The comparison groups strengthen the case: children who saw no model and children who watched an adult play quietly averaged about two such acts or fewer. These were not simply what nursery school children do with a mallet and a doll; they reached 25.8 only where they had been demonstrated. Boys imitated physical aggression most after watching a male model, which fits the idea that a model's characteristics affect how much is imitated. The evidence has a boundary: what was copied was aggression toward an inflatable toy, so the study supports observational learning as a route by which aggressive acts are acquired rather than proving it explains violence against people.

Where the points are earned
  • A: Research method (1): names the method, "an experiment", and supports the label with the manipulation of the model's behavior across three conditions. It does not stop at describing the one-way mirror, which would be procedure rather than design.
  • B: Operational definition (1): gives the measure as it was actually taken: coded five-second intervals, 240 of them, scored only when the child reproduced one of the model's specific acts. A general description of "aggressive behaviour" would not earn this.
  • C: Interpreting the statistic (1): interprets the mean difference in this study: about 25.8 against about two, a more than tenfold gap between groups matched on prior aggressiveness, and says what that licenses: the behavior came from what was watched, not from who the children were.
  • D: Ethical guideline (1): identifies parental consent, grounded in the sentence of the source that reports it, and adds protection from harm as a second guideline. Either would earn the point on its own.
  • E: Generalizability (1): makes a judgment and supports it with participant evidence from the source: 72 children, ages three to six, one university nursery school in the United States. It then separates the general principle from these particular numbers.
  • F: Argumentation, evidence (1 of 2): cites specific results: the distinctive modeled acts reproduced, about 25.8 imitative acts against about two without a model, imitation highest for boys watching a male model.
  • F: Argumentation, explanation (2 of 2): uses observational learning correctly, explaining that behavior was acquired by watching, without reinforcement and in the model's absence, and uses the control comparison to rule out the obvious rival explanation. Naming the boundary (a doll, not a person) shows control of the concept rather than weakening the argument.

AAQ 3 · 25 minutes · 7 points

Using the source, respond to parts A through F: do the reported results support or refute normative social influence as the explanation of conformity?

Directions

Read the source, then respond to all six parts in complete sentences, in order, labelling each part with its letter. You may write the parts as separate short paragraphs. Do not bring in a different study; every claim you make about the research must come from this source. Parts A through E are worth one point each and part F is worth two, for seven points in twenty-five minutes.

  • A. Identify the research method used in the study.
  • B. State the operational definition of the dependent variable, conformity.
  • C. Describe what the figure of about 37 percent of critical trials indicates about the study's findings.
  • D. Identify at least one ethical guideline the researchers applied.
  • E. Explain the extent to which the findings can be generalized beyond the people who took part, using specific evidence about the participants.
  • F. Explain how the study's specific results support or refute normative social influence as the explanation of conformity. (2 points)
Source

Source: research summary written for this course, based on Asch (1951, 1956).

The researchers examined how often a person will give an answer he can see is wrong when everyone else in the room has already given that answer aloud.

Across this series of studies 123 male college students in the United States took part, one at a time, in what each was told was an experiment on visual perception. Each participant joined a group of seven to nine other people, all of them confederates working with the experimenter: a fact the participant did not know. The group sat in a row, and the participant was seated second from the end so that he heard almost all of the other answers before giving his own.

On each of 18 trials the experimenter displayed two cards. One carried a single standard line; the other carried three comparison lines, one matching the standard in length and two differing from it clearly. Group members said aloud, in seating order, which comparison line matched the standard. On six of the trials the confederates gave the correct answer. On the remaining 12, the critical trials, every confederate named the same wrong line before the participant spoke.

Conformity was recorded as the proportion of critical trials on which the participant named aloud the same wrong line as the unanimous majority. Across participants that figure came to about 37 percent of critical responses. About 75 percent of participants went along with the majority on at least one critical trial and about a quarter never did. A control group judged the same cards alone and wrote their answers privately; those participants were wrong on under 1 percent of trials. In a further variation one confederate broke the unanimity by naming the correct line, and conformity fell to a fraction of its level in the unanimous condition.

Participants were interviewed after the session. Those who had gone along most often said that they had seen the difference between the lines but did not want to appear different from the group or to draw its attention; a smaller number said the group's agreement had made them doubt their own eyes.

Participation was voluntary and the line judgments carried no risk of harm. Participants were deceived about the purpose of the sessions and about the other people in the room, because telling them in advance would have destroyed the situation being studied. The interview afterward was used to explain the deception, the role of the confederates and the real purpose of the research before the participant left.

Your response
What a reader looks for on this prompt
  • A. The method, named: experiment (a laboratory experiment). The researchers arranged what the confederates said and compared it with a control condition in which people judged alone, so nothing here was left to occur naturally. "Naturalistic observation" and "case study" both contradict the source.
  • B. How the variable was counted here: the proportion of the 12 critical trials on which the participant said aloud the same wrong comparison line the majority had named. "How much the participant agreed with the group" restates the construct; the point comes from naming the critical trials and the spoken wrong answer.
  • C. Interpret the figure, and read it correctly: 37 percent is the share of critical responses that followed the majority, not the share of people who conformed: that is the 75 percent figure, and confusing the two is the commonest error on this item. Interpret it against the baseline: people judging alone erred on under 1 percent of trials, so more than a third of public answers went wrong under group pressure while nearly two thirds stayed independent.
  • D. A guideline they followed: debriefing (the interview afterward explained the deception, the confederates and the real purpose), voluntary participation, or protection from harm (a line-judging task carries no risk). Writing that the deception was unethical names a problem rather than a guideline applied, and earns nothing.
  • E. Participant evidence first, then the judgment: 123 male college students in the United States, tested in the 1950s, judging lines whose correct answer was obvious. The all-male, single-country, single-era sample is the evidence part E wants; cite it, then say how far the findings travel.
  • F. Results plus the concept, both (2 points): the strongest evidence for normative influence is the combination of the under 1 percent error rate alone, the interview reports of seeing the difference but not wanting to stand out, and the collapse of conformity when one confederate dissented. A full-credit answer also uses the minority who said they doubted their own eyes to limit the claim. One point for results without the concept, or the concept without results.
Show a 7/7 response

A The researchers used an experiment. They controlled what the confederates said on each trial and compared the result against a control condition in which people judged the same cards alone, so this is a laboratory experiment rather than observation of groups in natural settings.

B Conformity was operationally defined as the proportion of the 12 critical trials on which the participant said aloud the same wrong comparison line that the unanimous majority of confederates had named before him.

C About 37 percent means that just over a third of the answers given on critical trials followed the majority into an error. It is a rate of responses, not a count of people, so it does not mean 37 percent of participants conformed. Set against the control group, who were wrong on under 1 percent of trials when they judged alone, the figure indicates that hearing a unanimous wrong answer made public errors far more likely on a task whose answer was obvious, while leaving nearly two thirds of responses independent.

D The researchers debriefed their participants. The source reports that the interview after the session explained the deception, the role of the confederates and the real purpose of the research before the participant left. Participation was also voluntary, and the line-judging task itself carried no risk of harm.

E Generalization is limited by who was studied. All 123 participants were male college students in the United States, tested in the 1950s, so the results describe young American men of one decade and say nothing directly about women, older adults, or cultures where disagreeing carries a different cost. The task was also unusual: an unambiguous perceptual judgment made aloud among strangers. Conformity to a unanimous majority has been found in many places since, so the direction travels, but the 37 percent rate is not a fixed value.

F The results support normative social influence as the main explanation, though not the only one. Normative influence means going along with a group to gain acceptance or avoid standing out, not because you believe it is right. Three results point that way. First, control participants judging alone were wrong on under 1 percent of trials, so conformers could see which line matched; their errors were not failures of perception. Second, in the interviews those who had gone along most often said they had seen the difference but did not want to appear different from the group: a concern about being accepted, not about being correct, which is exactly what normative influence describes. Third, conformity fell sharply when a single confederate named the correct line: what changed was not the information about the lines but the cost of being the only person to disagree. The claim is qualified by the smaller group who reported that the majority's agreement made them doubt their own eyes. Those participants describe informational social influence, so the study supports normative influence as the dominant process here while refuting the stronger claim that it is all of conformity.

Where the points are earned
  • A: Research method (1): names the method, "an experiment", and backs the label with the manipulation of the confederates' answers and the alone condition used for comparison, instead of describing the line task and stopping.
  • B: Operational definition (1): gives the measure as it was taken: the proportion of the 12 critical trials on which the participant spoke the majority's wrong line aloud. It specifies the trials that counted and the behavior that was scored.
  • C: Interpreting the statistic (1): interprets 37 percent in context and corrects the usual misreading by stating that it is a rate of responses rather than a percentage of people, then anchors it to the under 1 percent baseline from the alone condition.
  • D: Ethical guideline (1): identifies debriefing and grounds it in what the source reports about the post-session interview, adding voluntary participation and a harmless task. Any one of the three would earn the point.
  • E: Generalizability (1): supports its judgment with participant evidence (123 male college students, United States, 1950s), and names the unusual feature of the task, then separates the general phenomenon from the particular rate.
  • F: Argumentation, evidence (1 of 2): cites three specific results: the under 1 percent error rate in the alone condition, the interview reports, and the drop in conformity when one confederate dissented.
  • F: Argumentation, explanation (2 of 2): defines normative social influence correctly and links each result to it: accurate private perception, a stated motive of acceptance rather than accuracy, and unanimity rather than information doing the work. Using the minority who doubted their own eyes to mark where informational influence takes over shows the concept is under control rather than merely named.

AAQ 4 · 25 minutes · 7 points

Using the source, respond to parts A through F. Explain whether the findings support or refute neuroplasticity: the claim that experience changes the physical structure of the brain.

Directions

You have 25 minutes. Respond to all parts of the question in complete sentences, in order, labelling each part A through F. Use the source to answer; do not summarize it.

  • A. Identify the research method used in the study.
  • B. State the operational definition of the dependent variable, cortical development.
  • C. Describe what the difference in mean cortical weight between the enriched and isolated groups indicates about the study's findings.
  • D. Identify at least one ethical guideline the researchers applied.
  • E. Explain the extent to which the findings can be generalized, using specific evidence about the participants.
  • F. Explain how the study's specific results support or refute neuroplasticity. (2 points)
Source

Source: research summary written for this course, based on Rosenzweig, Bennett, and Diamond (1972).

Researchers at a university laboratory asked whether the complexity of an animal's daily surroundings changes the physical structure of its brain. The animals were male rats of a standard laboratory strain, taken from the same litters and weaned at about twenty-five days old. Within each litter, siblings were assigned by chance to one of three housing conditions, so that each litter contributed one animal to each condition and inherited differences were spread evenly across the groups.

In the enriched condition, about a dozen rats shared one large cage furnished with a changing set of objects (ladders, platforms, wheels, boxes, tunnels) drawn from a larger pool, with several objects swapped for different ones each day. In the standard colony condition, two or three rats lived together in an ordinary laboratory cage containing no objects. In the isolated condition, a rat lived alone in a cage of the same kind, in a separate quiet room. Food and water were freely available in every condition, and lighting, temperature, diet and handling were held the same across the three. Animals remained in the assigned condition for about thirty days in the standard version of the experiment.

At the end of that period the brains were removed and dissected into standard sections. Cortical development was recorded as the weight in milligrams of the dissected cortical tissue, the thickness of the cortex measured in prepared sections, and the ratio of cortical weight to subcortical weight, which the researchers treated as their most consistent measure because it adjusts for differences in overall body and brain size. They also measured the activity of acetylcholinesterase, an enzyme involved in synaptic transmission.

Rats from the enriched condition finished with a heavier and thicker cortex than their isolated littermates. Averaged across the cortex as a whole, the difference in mean weight was about four percent; in the occipital region, the visual area at the back of the cortex, it was larger, on the order of six percent. The cortical-to-subcortical ratio separated the groups more consistently than raw weight did, and total acetylcholinesterase activity was greater in the cortex of the enriched animals. The differences were small in absolute terms, but they appeared again in repeated runs of the experiment over several years, and standard colony animals fell between the other two groups.

The work was carried out under the institution's animal-care oversight. Housing met the standards then in force for cage size, ventilation, temperature, diet and veterinary attention, trained staff handled the animals daily, and the procedures at the end of the study followed approved humane methods.

Your response
What a reader looks for on this prompt
  • A: Research method: the word experiment. Littermates were assigned to housing conditions by chance and the researchers controlled the condition each animal lived in, so housing is a manipulated independent variable. Saying "the researchers put rats in different cages and weighed their brains" describes the procedure without naming the method and earns nothing; so does calling it a correlational study or a naturalistic observation.
  • B: Operational definition: cortical development as this study measured it: the weight in milligrams of dissected cortical tissue, cortical thickness in prepared sections, and the ratio of cortical to subcortical weight. "Brain development" or "how developed the brain was" restates the construct rather than defining it by the measurement, and earns nothing.
  • C: Interpreting the statistic: say what the roughly four percent difference in mean cortical weight tells us here: on average the enriched rats ended with a measurably heavier cortex than their isolated littermates, a small difference that ran in the same direction across replications and was larger, about six percent, in the occipital cortex. Defining what a mean or a percentage is does not earn the point.
  • D: Ethical guideline: institutional animal-care review, or humane treatment and housing of animal subjects: cage standards, ventilation, temperature, diet, veterinary attention, trained handling, approved humane procedures at the end of the study. Point to what the source reports. Criticizing the isolation condition names a problem, not a guideline the researchers applied, and earns nothing.
  • E: Generalizability: limited, and the participant evidence has to be specific: male rats of one laboratory strain, littermates weaned at about twenty-five days, housed for roughly thirty days in a laboratory, with no human participants and no females. The contrast between about a dozen cage-mates with daily new objects and solitary housing in a quiet room is also far wider than the gap between two children's homes. A bare "this can't be generalized to humans" with no participant detail earns nothing.
  • F: Argumentation (2 points): commit to support, then quote the study's own numbers and connect them to the concept. Neuroplasticity claims that experience changes brain structure; the enriched animals' cortex was about four percent heavier, about six percent heavier in the occipital region, thicker, higher in cortical-to-subcortical ratio and higher in total acetylcholinesterase activity, and random assignment of littermates rules out inherited differences as the cause. Results with no link to the concept, or the concept with no results, caps the part at 1.
Show a 7/7 response

A The researchers used an experiment. Littermate rats were assigned by chance to one of three housing conditions, and the researchers controlled which condition each animal lived in, so housing was a manipulated independent variable rather than something the animals arrived with.

B The dependent variable was cortical development, and in this study it was operationally defined as the weight in milligrams of the dissected cortical tissue, the thickness of the cortex measured in prepared sections, and the ratio of cortical weight to subcortical weight, the measure they relied on most because it adjusts for differences in overall brain and body size.

C The statistic is the roughly four percent difference in mean cortical weight between the enriched and the isolated groups. It indicates that the average rat raised with cage-mates and rotating objects finished with a measurably heavier cortex than the average littermate raised alone. Four percent is a small difference in absolute terms, but it is a difference between group averages that came out in the same direction each time the experiment was repeated, and it was larger, about six percent, in the occipital cortex.

D One ethical guideline the researchers applied was institutional animal-care review with humane treatment of animal subjects. The source reports that the work was carried out under institutional animal-care oversight, that housing met the standards then in force for cage size, ventilation, temperature, diet and veterinary attention, and that the procedures at the end of the study followed approved humane methods.

E The findings generalize only so far. The participants were male rats of a single standard laboratory strain, littermates weaned at about twenty-five days old and kept in their assigned cages for roughly thirty days in a laboratory. No humans took part, no females were included, and only one strain was used, so the size of the effect cannot be carried over to people. The conditions were also more extreme than ordinary human variation: about a dozen cage-mates with new objects daily at one end, a rat living alone in a quiet room at the other. What does travel is the claim the study was built to test: that structure responds to experience at all.

F The results support neuroplasticity, the claim that experience changes the physical structure of the brain. The enriched rats' cortex was about four percent heavier on average than their isolated littermates', about six percent heavier in the occipital region, thicker in prepared sections, higher in the ratio of cortical to subcortical weight, and higher in total acetylcholinesterase activity. F Those are structural and biochemical differences that line up with the one thing that differed between the groups: what the animals did all day. Because littermates were assigned to conditions by chance, inherited differences cannot explain the gap, so experience is the credible cause of the structural change, which is what neuroplasticity claims. Standard colony animals landing between the other two groups strengthens the reading, since the amount of structural change tracked the amount of experience.

Where the points are earned
  • A: Research method (1): the first sentence names the method, and the rest of A justifies the name by pointing to random assignment of littermates and to a manipulated independent variable. Naming the method is what the rubric requires; describing the procedure alone would earn nothing.
  • B: Operational definition (1): B gives the measurements this study actually took (milligrams of dissected cortical tissue, thickness in prepared sections, and the cortical-to-subcortical weight ratio) instead of defining cortical development in the abstract.
  • C: Interpreting the statistic (1): C says what the four percent mean difference indicates in this study: the average enriched rat finished with a heavier cortex than the average isolated littermate, a small but repeated difference that was larger in the occipital cortex. It interprets the number rather than defining what a percentage is.
  • D: Ethical guideline (1): D names institutional animal-care review and humane housing, both grounded in what the source reports, and is a guideline the researchers followed rather than a criticism of the isolation condition.
  • E: Generalizability (1): E gives a judgment, limited, and supports it with specific participant evidence: male rats, one laboratory strain, littermates weaned at about twenty-five days, thirty days of housing, a laboratory setting, and conditions far more extreme than ordinary human variation.
  • F: Argumentation (2): the first point comes from citing the study's specific results: the four percent whole-cortex difference, the six percent occipital difference, greater thickness, a higher cortical-to-subcortical ratio, greater acetylcholinesterase activity. The second point comes from the explanation that follows: these are structural differences produced by the one factor that varied, random assignment of littermates rules out heredity, and the standard colony group falling in between shows the change scaling with experience, which is precisely the claim neuroplasticity makes.

AAQ 5 · 25 minutes · 7 points

Using the source, respond to parts A through F. Explain whether the findings support or refute the cognitive explanation of depression: the claim that changing distorted thinking changes mood.

Directions

You have 25 minutes. Respond to all parts of the question in complete sentences, in order, labelling each part A through F. Use the source to answer; do not summarize it.

  • A. Identify the research method used in the study.
  • B. State the operational definition of the dependent variable, depression severity.
  • C. Describe what the difference between the two groups' mean end-of-treatment scores on the self-report depression inventory indicates about the study's findings.
  • D. Identify at least one ethical guideline the researchers applied.
  • E. Explain the extent to which the findings can be generalized, using specific evidence about the participants.
  • F. Explain how the study's specific results support or refute the cognitive explanation of depression. (2 points)
Source

Source: research summary written for this course, based on Rush, Beck, Kovacs, and Hollon (1977).

Researchers at the outpatient clinic of a university medical center in the United States compared a talking therapy with an antidepressant medication as treatments for depression. Forty-one adults who had come to the clinic for help with moderate to severe depression took part. Women and men were included; all were seeking treatment for themselves, and each was interviewed and rated by a clinician before entering. People whose difficulties were better accounted for by another condition, and people already receiving another treatment, were not enrolled.

Each patient was assigned at random to one of two twelve-week treatments. Nineteen were assigned to cognitive therapy: up to twenty sessions in which the therapist and the patient identified automatic negative thoughts about the self, the world and the future, tested those thoughts against evidence from the patient's own week, and worked out more accurate alternatives, with assignments to carry out between sessions. Twenty-two were assigned to imipramine, a tricyclic antidepressant, prescribed and adjusted by a physician across the same twelve weeks and tapered before the end of the trial.

Depression severity was measured with two standardized instruments given before treatment, at intervals during it, and at the end: an inventory the patient completed alone and a rating scale completed by a clinician who interviewed the patient. Each instrument yields a total score, with higher totals indicating more severe symptoms.

Both groups improved over the twelve weeks. Mean self-report scores at intake were close to 30 in both groups, in the moderate to severe range. By the end of treatment the cognitive therapy group's mean had fallen into the single digits and the medication group's mean to the mid-teens, leaving a difference between the two group means of roughly seven points in favor of cognitive therapy. The clinician ratings moved in the same direction. About 79 percent of the patients assigned to cognitive therapy were rated markedly improved or free of symptoms at the end of treatment, compared with about 23 percent of those assigned to medication. One of the nineteen patients in cognitive therapy left treatment early, compared with eight of the twenty-two taking medication. The trial ran at one clinic, with no placebo condition and no group receiving both treatments.

Before entering, each patient was told what the two treatments involved, what would be measured and how often, and that participation could be ended at any point without affecting the care available at the clinic; each then gave informed consent. A clinician monitored symptoms at scheduled points throughout the twelve weeks. Patients whose symptoms did not improve, and patients who left the trial early, were offered continued care and referral to other treatment. Records were kept confidential.

Your response
What a reader looks for on this prompt
  • A: Research method: the word experiment, or the more precise randomized clinical trial, which is an experiment. Patients were assigned at random to one of two treatments, so treatment type is a manipulated independent variable. Calling it a case study, a correlational study or a longitudinal study earns nothing, and so does describing the twelve weeks of treatment without naming the method.
  • B: Operational definition: depression severity as this study measured it: the total score on a standardized self-report inventory the patient completed and on a rating scale completed by an interviewing clinician, given before, during and after the twelve weeks, with higher totals meaning more severe symptoms. Defining depression by its symptoms is a definition of the construct, not an operational definition, and earns nothing.
  • C: Interpreting the statistic: say what the roughly seven-point gap between the two group means indicates here: both groups started near 30 and both improved, but the average patient finishing cognitive therapy ended with a lower symptom score than the average patient finishing on medication. Note that it compares group averages, so it does not mean every person in cognitive therapy did better. Defining what a mean difference is earns nothing.
  • D: Ethical guideline: informed consent is the cleanest answer: patients were told what each treatment involved and what would be measured before agreeing. The right to withdraw without penalty, protection from harm through scheduled clinical monitoring and the offer of continued care and referral, and confidentiality of records are all equally creditable. Naming the missing placebo condition is a design criticism, not a guideline applied, and earns nothing.
  • E: Generalizability: limited, with the participant evidence named: 41 adults, women and men, at a single university outpatient clinic in one country, all of them help-seekers who came to the clinic themselves with moderate to severe depression, and people with other conditions or already in treatment excluded. Forty-one people at one site is a small base for a treatment claim, and the findings say nothing about children or adolescents, about people with mild depression, or about people who never seek treatment. A bare "the sample was too small" with no participant detail earns nothing.
  • F: Argumentation (2 points): commit to support, then use the numbers. The cognitive explanation holds that depressed mood is maintained by distorted negative thinking about the self, the world and the future; a treatment that targets those thoughts directly should therefore reduce symptoms. It did (roughly seven points lower on the mean self-report score, about 79 percent versus about 23 percent rated markedly improved or free of symptoms, and 1 dropout out of 19 versus 8 out of 22), and random assignment makes the treatment the credible cause. A strong response also names the limit: symptoms were measured, thoughts were not, so the result is consistent with the explanation rather than proof of the mechanism. Results with no link to the concept, or the concept with no results, caps the part at 1.
Show a 7/7 response

A The researchers used an experiment: a randomized clinical trial. Patients were assigned at random to either cognitive therapy or imipramine, so treatment type was a manipulated independent variable rather than something patients chose.

B The dependent variable was depression severity, operationally defined as the total score on two standardized instruments given before, during and after the twelve weeks: an inventory the patient completed alone and a rating scale completed by a clinician who interviewed the patient, with higher totals meaning more severe symptoms.

C The statistic is the roughly seven-point difference between the two groups' mean self-report scores at the end of treatment. Both groups began near 30, in the moderate to severe range, and both improved, so the seven points indicate how much further the cognitive therapy group moved: the average patient finishing cognitive therapy ended in the single digits while the average patient finishing on medication ended in the mid-teens. It is a comparison of group averages, so it does not mean every patient in cognitive therapy did better than every patient on medication.

D One ethical guideline the researchers applied was informed consent. The source reports that before entering, each patient was told what the two treatments involved and what would be measured, and was told that participation could end at any point without affecting the care available at the clinic. Protection from harm was also handled: a clinician monitored symptoms at scheduled points, and patients who did not improve or who left early were offered continued care and referral.

E The findings generalize to patients much like these and not far beyond. There were 41 adults, women and men, all of them people who came on their own to one university outpatient clinic in one country with moderate to severe depression, and people with other conditions or already in treatment were excluded. That is a small sample at a single site, so the exact size of the advantage is imprecise. The results also say nothing about children or adolescents, about people with mild depression, or about people who never seek treatment.

F The results support the cognitive explanation of depression. Patients given a treatment aimed directly at distorted thinking ended about seven points lower on the mean self-report inventory than patients given medication, about 79 percent of them were rated markedly improved or free of symptoms against about 23 percent on medication, and only one of nineteen left treatment early against eight of twenty-two. F The cognitive explanation says depressed mood is maintained by negative, distorted beliefs about the self, the world and the future, so a therapy that catches those thoughts and tests them against evidence should lift mood, and here it did, at least as much as a drug acting on neurotransmitters, with random assignment making the treatment the credible cause. The support is real but partial: the study measured symptoms, not the thoughts themselves, so it shows the prediction holding rather than proving that changed thinking is the mechanism.

Where the points are earned
  • A: Research method (1): the first sentence names the method, experiment, and identifies the more precise form, a randomized clinical trial, then justifies the label with random assignment and a manipulated independent variable.
  • B: Operational definition (1): B gives the measurement this study used (total scores on a patient-completed inventory and a clinician-completed rating scale, taken before, during and after treatment, higher meaning more severe) instead of defining depression as a construct.
  • C: Interpreting the statistic (1): C says what the seven-point gap indicates in this study: both groups started near 30 and improved, and the cognitive therapy group finished further down, in the single digits against the mid-teens. The closing sentence shows the response understands that a difference in means describes groups, not individuals.
  • D: Ethical guideline (1): D names informed consent and grounds it in what the source reports, then adds protection from harm through monitoring and the offer of continued care. Both are guidelines the researchers followed, not criticisms of the design.
  • E: Generalizability (1): E gives a judgment and backs it with specific participant evidence: 41 adults, women and men, one university outpatient clinic in one country, self-referred help-seekers with moderate to severe depression, other conditions excluded, then says which populations are therefore out of reach.
  • F: Argumentation (2): the first point comes from citing the study's specific results: the seven-point difference in mean self-report scores, about 79 percent against about 23 percent rated markedly improved or free of symptoms, one dropout of nineteen against eight of twenty-two. The second point comes from the explanation that follows: the cognitive explanation is stated correctly, the prediction it makes is spelled out, the results are shown to meet that prediction with random assignment supporting the causal reading, and the honest limit (symptoms measured, thoughts not) keeps the claim the size the evidence allows.

EBQ 1 · 45 minutes · 7 points

Using the sources provided, develop and justify an argument about whether high schools should start the school day later in the morning.

Directions

You have 45 minutes. Use the three sources below. Your evidence must come from at least two of them. Cite the sources you use as "Source A," "Source B," or "Source C," or by the authors' names. Write in complete sentences and label each part.

  • A. Propose a specific and defensible claim, based in psychological science, that responds to the question.
  • B(i). Support your claim using one specific and relevant piece of evidence from one of the sources, and cite that source.
  • B(ii). Explain how the evidence in Part B(i) supports your claim, using a psychological perspective, theory, concept, or research finding you learned in this course.
  • C(i). Provide additional support for your claim using a second specific and relevant piece of evidence, drawn from a different source than the one used in Part B(i), and cite that source.
  • C(ii). Explain how the evidence in Part C(i) supports your claim, using a psychological perspective, theory, concept, or research finding that is different from the one used in Part B(ii).
Source A

Source A: research summary written for this course, based on Carskadon and colleagues (1980, 1998).

A sleep laboratory studied adolescents in a residential summer sleep camp, recording brain activity overnight and testing daytime alertness the next day. Participants were boys and girls ranging from late childhood into the late teens, some of them followed across several summers as they moved through puberty. Given a fixed ten-hour opportunity to sleep each night, the older adolescents slept about as long as the youngest, roughly nine hours, even though the older ones reported feeling less sleepy in the evening. Compared across stages of pubertal maturation, the onset of melatonin release shifted later, so the more physically mature adolescents did not become sleepy until later at night. A second study by the same group followed several dozen students in a Rhode Island district whose high school day began at 7:20 a.m. On school nights those students slept roughly seven hours, well short of the nine hours the laboratory work indicated they needed. On morning laboratory tests many fell asleep within a few minutes, a degree of sleepiness clinicians treat as pathological, and some entered REM sleep directly, a pattern otherwise seen in sleep disorders. Limitation: the camp work used small volunteer samples living under artificial conditions, and the school study compared students before and after a schedule change rather than assigning them to start times, so maturation and course load changed alongside the clock.

Source B

Source B: research summary written for this course, based on Walker and colleagues (2002).

Healthy young adults with no reported sleep disorders learned a motor sequence task in a laboratory: typing a fixed five-key sequence with the non-dominant hand, as quickly and accurately as possible, across a series of timed trials. Learning was operationalized as the number of correct sequences typed per trial. Participants were assigned to be retested either after about twelve hours awake across an ordinary day or after about twelve hours that included a night of sleep. No one practiced between training and retest. The group that stayed awake finished close to where training had left them. The group that slept typed the sequence roughly 20 percent faster and made fewer errors, with no additional practice. When the researchers examined the overnight sleep recordings, the size of a participant's improvement was related to the amount of stage 2 NREM sleep in the last quarter of the night, the part of the night a person loses by setting an early alarm. A companion experiment found that a daytime nap containing both slow-wave and REM sleep produced gains comparable to a full night, while shorter naps without REM did not. Limitation: the samples were small groups of healthy young adults and the task was narrow, so a keyboard sequence is not the same as classroom material.

Source C

Source C: research summary written for this course, based on Dunster and colleagues (2018).

Two public high schools in Seattle moved their start time from 7:50 a.m. to 8:45 a.m., a delay of 55 minutes, and researchers measured what changed. Participants were roughly 170 sophomore biology students, studied as two cohorts: one in the spring before the change and one in the spring after it. Each student wore a wrist activity monitor for about two weeks, which recorded movement and rest and produced an objective estimate of sleep duration rather than a self-report. Students also completed sleepiness questionnaires, and the researchers obtained course grades and attendance records. After the change, median sleep duration on school nights was about 34 minutes longer. Self-reported daytime sleepiness was lower. Median grades in the same course, taught with the same curriculum both years, were about 4.5 percent higher in the later-start cohort, and at the school serving more economically disadvantaged students, tardiness and first-period absence dropped. Limitation: this is a before-and-after comparison of two different groups of students, not a randomized experiment, so anything else that differed between the two years (the teachers, the particular students, the season) could contribute to the grade difference, and the comparison covers one course rather than a whole transcript.

Your response
What a reader looks for on this prompt
  • A. The claim (1 point). Take a side and make it specific enough to be wrong. "Schools should start later because adolescent sleep timing is biologically delayed and the lost sleep is the sleep that consolidates learning" earns the point. "Sleep is important for students" does not: nothing in the sources could contradict it. "Later starts are not worth the cost because the grade gain in Source C comes from an uncontrolled before-and-after comparison" is also defensible here, and the sources support it; you are not required to argue for delay.
  • B(i) and C(i): the evidence (1 point each). Quote a particular result, not a topic. Usable pieces: the roughly nine hours of sleep obtained under a ten-hour opportunity and the later melatonin onset with pubertal maturation (Source A); the seven hours slept under a 7:20 a.m. start and falling asleep within minutes on morning tests (Source A); the roughly 20 percent overnight gain with no extra practice, and its link to stage 2 NREM sleep late in the night (Source B); the 55-minute delay, the 34 additional minutes of measured sleep, the 4.5 percent higher median grades, or the drop in first-period tardiness (Source C). "Source C is about Seattle schools" earns nothing. The two pieces must come from two different sources: two findings out of Source A cannot earn both points.
  • B(ii). The first explanation (2 points). Name the concept and apply it. Circadian rhythms and the suprachiasmatic nucleus is the cleanest fit for Source A: the SCN sets a roughly 24-hour rhythm from light and drives melatonin, the adolescent phase delay pushes sleep onset later, and therefore an early bell cannot move bedtime; it only cuts the morning off. One point instead of two if you write "teenagers' body clocks are different" without naming the rhythm, the nucleus, or the phase delay, or if you only restate that the students slept seven hours.
  • C(ii). The second explanation (2 points), with a different concept. Sleep-dependent memory consolidation is the natural second family: the hippocampus-to-cortex stabilization that happens during NREM and REM explains why 34 more minutes of sleep shows up as higher grades, and Source B supplies the research finding that the relevant sleep sits in the last quarter of the night. Other usable concepts: operant conditioning and environmental contingencies for the attendance and tardiness result, or selective attention and working memory capacity for daytime performance.
  • Which pairs count as "different." Circadian rhythm and the SCN (a biological explanation of timing) paired with memory consolidation (a cognitive explanation of learning) satisfies the rule, as does circadian rhythm paired with operant reinforcement of attendance, or consolidation paired with the person-environment fit between a delayed clock and a fixed bell. What fails: "circadian rhythms" and "the body clock," or "sleep deprivation" and "not enough sleep." Those are one concept named twice, and the second explanation scores 1 at most.
Show a 7/7 response

A High schools should start the day later, because the obstacle to adolescent sleep is biological timing rather than teenage choice, and the sleep an early bell removes is precisely the sleep that stabilizes what students learned the day before.

B(i) Source A reports that adolescents given a ten-hour sleep opportunity slept about nine hours regardless of age, that the onset of melatonin release shifted later as students matured through puberty, and that students in a district starting at 7:20 a.m. slept only about seven hours on school nights and fell asleep within a few minutes on morning laboratory tests.

B(ii) This supports the claim through circadian rhythms and the suprachiasmatic nucleus. The SCN uses light to hold the body on a roughly 24-hour cycle and triggers melatonin release when that cycle says night has arrived, and Source A shows that the release point moves later with pubertal maturation. That matters because a start time can only change when sleep ends, not when it begins: a student whose melatonin does not rise until late cannot fall asleep on command at nine o'clock, so moving the bell from 8:45 to 7:20 subtracts time from the front of the morning rather than adding it to the night. The seven-hour average and the pathological-level sleep latencies in Source A are what that subtraction looks like. The shortfall is a mismatch between a fixed schedule and a delayed biological clock, which is a problem schools can fix and adolescents cannot.

C(i) Source C provides a second kind of support: when two Seattle high schools delayed their start by 55 minutes, wrist monitors showed students sleeping about 34 minutes longer on school nights, and median grades in the same biology course were about 4.5 percent higher in the later-start cohort.

C(ii) Sleep-dependent memory consolidation explains why extra sleep turns into higher grades. Consolidation is the process by which a newly encoded memory is stabilized and reorganized during sleep rather than during practice, and Source B demonstrates it directly: participants who slept typed a learned sequence about 20 percent faster with no further practice, while those who stayed awake did not improve, and the size of each person's gain tracked the stage 2 NREM sleep in the final quarter of the night. The 34 minutes Seattle students recovered came from exactly that final quarter, so the later start returned the sleep in which the previous day's learning gets consolidated. The grade increase is the classroom version of the laboratory result. I would hold this conclusion loosely, since Source C compared two different cohorts rather than randomly assigning start times, but the mechanism in Source B is experimental and predicts the direction Source C found.

Where the points are earned
  • A: Defensible claim (1): The claim answers the question asked with a position, not a summary, and names the reasoning the response then defends: the constraint is biological timing, and the sleep lost is the sleep that consolidates learning. Sources could contradict it, if Source A had shown melatonin onset unchanged by puberty, the claim would fail.
  • B(i): Evidence from a source (1): Three specific results from Source A, cited by source: nine hours slept under a ten-hour opportunity, later melatonin onset with pubertal maturation, and seven hours of school-night sleep with sleep latencies of a few minutes under a 7:20 a.m. start. Each is a particular finding rather than a description of what the source is about.
  • B(ii): Explanation with a concept (2): Names circadian rhythms and the suprachiasmatic nucleus and applies them correctly: the SCN entrains to light and gates melatonin, the adolescent phase delay moves sleep onset later, and therefore an earlier start clips the end of sleep instead of moving the beginning. The explanation connects the mechanism to the exact numbers cited rather than restating them.
  • C(i): Evidence from a different source (1): The second piece comes from Source C, not Source A, and is equally specific: a 55-minute delay, about 34 additional minutes of actigraphy-measured sleep, and median grades about 4.5 percent higher in the same course.
  • C(ii): Explanation with a different concept (2): Sleep-dependent memory consolidation is a genuinely different family of explanation from circadian timing (one is about when sleep happens, the other about what sleep does to a memory), and it is named, defined, and applied. It also imports Source B's research finding, the 20 percent overnight gain tied to late-night stage 2 NREM sleep, to link the recovered 34 minutes to the grade increase. The closing sentence acknowledges Source C's cohort design without abandoning the claim, which costs nothing and shows the reader the argument is controlled.

EBQ 2 · 45 minutes · 7 points

Using the sources provided, develop and justify an argument about the extent to which the situation, rather than personality, explains harmful obedience.

Directions

You have 45 minutes. Use the three sources below. Your evidence must come from at least two of them. Cite the sources you use as "Source A," "Source B," or "Source C," or by the authors' names. Write in complete sentences and label each part.

  • A. Propose a specific and defensible claim, based in psychological science, that responds to the question.
  • B(i). Support your claim using one specific and relevant piece of evidence from one of the sources, and cite that source.
  • B(ii). Explain how the evidence in Part B(i) supports your claim, using a psychological perspective, theory, concept, or research finding you learned in this course.
  • C(i). Provide additional support for your claim using a second specific and relevant piece of evidence, drawn from a different source than the one used in Part B(i), and cite that source.
  • C(ii). Explain how the evidence in Part C(i) supports your claim, using a psychological perspective, theory, concept, or research finding that is different from the one used in Part B(ii).
Source A

Source A: research summary written for this course, based on Milgram (1963, 1974).

Forty men from the New Haven area, aged roughly 20 to 50 and recruited by newspaper advertisement, were told the study concerned learning and memory. Each was assigned the role of teacher at a shock generator whose switches were labelled from 15 to 450 volts. When the learner, who was a confederate, answered incorrectly, an experimenter in a lab coat instructed the teacher to move one switch higher and delivered a fixed series of verbal prods whenever the teacher objected. No shocks were actually delivered. The dependent variable was the highest switch a participant pressed. Twenty-six of the forty men, 65 percent, continued to the maximum level, and many showed visible distress while continuing. The same researcher later ran variations that changed only the setting. When the experimenter left the room and gave the orders by telephone, the proportion reaching the maximum fell to about one participant in five. When the study moved from the university to a run-down office building, it fell to roughly half. When two confederates playing fellow teachers refused and walked out first, about one in ten continued to the end. Limitation: all participants were men from one American city, the deception and distress would not pass review today, and no participant was randomly assigned to a no-authority control.

Source B

Source B: research summary written for this course, based on Burger (2009).

A partial replication was designed to run under modern ethical rules. Seventy adults, 29 men and 41 women aged 20 to 81, were recruited from a California community. A two-stage screening excluded anyone with a history of anxiety or depressive disorders and anyone who recognized the original procedure, a clinical psychologist observed every session with authority to halt it, participants were told at least three times that they could stop and keep the payment, and the sample shock was reduced. The procedure was stopped at 150 volts, the point in the original study where the learner first demands to be released and where most participants who defied the experimenter did so. About 70 percent of participants in the base condition were still going at that point and had to be stopped by the experimenter, a rate close to the comparable figure in the original study; the difference was not statistically significant. Men and women did not differ. Participants had completed questionnaire measures of empathic concern and desire for control weeks earlier, and those scores showed no reliable difference between the people who continued and the people who refused. Limitation: the procedure stopped at 150 volts, so the study cannot show what any participant would have done at higher levels, and the screening produced an unusually resilient sample.

Source C

Source C: research summary written for this course, based on Haslam and Reicher (2006).

Fifteen men, selected from a large applicant pool after screening for mental health and antisocial tendencies, were randomly assigned to be five guards or ten prisoners in a purpose-built institution and observed for eight days. Clinical psychologists monitored participants daily, an independent ethics panel held the power to stop the study at any moment, and participants could withdraw. Psychological measures, including social identification with one's own group, were collected repeatedly. The guards never became a cohesive group: they disagreed about how to use their authority, and their identification with the guard role declined across the days. The prisoners moved the other way. Once it became clear that no prisoner could be promoted into the guard group, their identification with each other rose sharply, and on day six they breached the barrier and the guard system collapsed. The participants then set up a self-governing commune, which broke down within a day, after which a subgroup began planning a far harsher regime and the researchers ended the study. The authors argued that people do not automatically conform to an assigned role; harm follows from identifying with a group and its leaders and working actively toward what feels like a worthwhile cause. Limitation: fifteen male volunteers, filmed for broadcast, is a small and observed sample.

Your response
What a reader looks for on this prompt
  • A. The claim (1 point). The question asks about extent, so a claim that weighs the two causes earns the point: "situational features explain most of the variation in obedience, because identical people obey at very different rates when the setting changes, while measured personality differences do not separate those who obey from those who refuse." A qualified claim scores equally well, "the situation explains most of it, but Source C shows the situational cause is identification with a group and its leaders, not role-following", and is often easier to defend. "Both personality and the situation matter" does not earn the point; nothing in the sources could contradict it.
  • B(i) and C(i): the evidence (1 point each). Name a result. Usable pieces: 26 of 40 men, 65 percent, reaching the maximum level, or the drops to about one in five with telephoned orders and about one in ten after two confederates walked out (Source A); the roughly 70 percent still going at 150 volts under modern safeguards, or the empathic-concern and desire-for-control scores failing to separate those who continued from those who stopped (Source B); guards' identification with their role falling while prisoners' identification with each other rose, the collapse of the guard system on day six, or the later push toward a harsher regime (Source C). The two pieces must come from two different sources.
  • B(ii). The first explanation (2 points). The agentic state and diffusion of responsibility fit Source A precisely: a legitimate authority who is physically present and who accepts responsibility lets a person stop seeing themselves as the author of their own act, so weakening any of those cues (moving the authority to a telephone, removing the university's prestige, or breaking the unanimity of compliance) lowers obedience without changing who the participants are. Normative and informational social influence work equally well for the rebellious-peers variation. One point rather than two if the response says "the situation made them do it" without naming a mechanism, or names the agentic state but never ties it to a particular variation.
  • C(ii). The second explanation (2 points), with a different concept. Trait theory and the person-situation controversy fit Source B: traits predict behavior averaged over many occasions, while the situation usually decides any single occasion, which is exactly the pattern when empathy and control scores fail to predict a one-time decision. Social identity theory fits Source C: guards who never formed a shared identity could not act as a group, while prisoners who did, acted decisively.
  • Which pairs count as "different." Agentic state and diffusion of responsibility (a situational account of how responsibility is displaced) paired with trait theory and the person-situation debate (a dispositional account of prediction) satisfies the rule, as does agentic state paired with social identity theory, or normative social influence paired with social identity theory. What fails: "obedience" and "social influence," or "conformity" and "going along with the group." Those are one idea named twice and cap the second explanation at 1 point.
Show a 7/7 response

A The situation carries most of the explanatory weight: obedience rates swing enormously when only the setting is changed, while measured personality differences do not separate the people who obey from the people who refuse. The qualification Source C adds is that the situational cause is identification with a group and its leaders rather than automatic role-following.

B(i) Source A reports that 26 of 40 men, 65 percent, continued to the maximum switch in the original procedure, but that the proportion fell to about one participant in five when the experimenter gave orders by telephone instead of standing in the room, and to about one in ten when two confederates playing fellow teachers refused and walked out first.

B(ii) The agentic state, together with diffusion of responsibility, explains this. Milgram's account is that a person facing a legitimate authority shifts out of an autonomous state, in which they see themselves as the author of their own actions, into an agentic state, in which they see themselves as carrying out someone else's intention and hand the responsibility for the outcome to that person. Each variation removes one of the cues that supports the shift. A voice on a telephone is a weaker claim to authority than a man in a lab coat two feet away, so the participant is left holding the responsibility himself. Two peers who walk out destroy the unanimity that made compliance look like the only available reading of the situation. The participants in those variations were drawn from the same population as the original sample, so nothing about their personalities changed between conditions. What changed was the strength of the situational cues, and behavior followed.

C(i) Source B adds evidence of the other kind. In a replication run under modern safeguards, questionnaire measures of empathic concern and desire for control, completed weeks before the session, showed no reliable difference between the participants who continued past 150 volts and those who refused, even though about 70 percent continued.

C(ii) That null result is what trait theory and the person-situation controversy predict. Traits are stable dispositions that show their predictive power when behavior is averaged across many occasions; a single decision in a strong, unfamiliar situation is precisely the case where the situation dominates and trait scores explain almost nothing. Empathic concern probably does predict how often a person is kind over a year, but it does not predict what one person does once at a shock generator with an experimenter insisting. Because Source B measured personality directly and prospectively rather than inferring it after the fact, the absence of a relationship is informative: if harmful obedience were a property of unusual people, those questionnaires were the place it should have appeared.

Where the points are earned
  • A: Defensible claim (1): The claim answers the "extent" question with a weighting rather than a summary, gives the reason the response then defends, and is falsifiable: had Source B found empathic concern predicting refusal, the claim would be in trouble. The qualification drawn from Source C sharpens it instead of hedging it away.
  • B(i): Evidence from a source (1): Specific and accurate figures from Source A, cited by source, 26 of 40 at the maximum level, roughly one in five with telephoned orders, roughly one in ten after two confederates refused. The comparison across variations is what makes the evidence relevant to a claim about situations.
  • B(ii): Explanation with a concept (2): Names the agentic state and diffusion of responsibility, defines the autonomous-to-agentic shift correctly, and maps each variation onto a cue that the shift depends on. It also states the logical point that holds the argument together: the participants did not change across conditions, only the setting did.
  • C(i): Evidence from a different source (1): The second piece comes from Source B rather than Source A, and is a particular result: prospectively measured empathic concern and desire for control failing to distinguish those who continued from those who refused, alongside the roughly 70 percent continuation rate.
  • C(ii): Explanation with a different concept (2): Trait theory and the person-situation controversy are a different family of explanation from the agentic state: one is a dispositional theory about when traits predict behavior, the other a situational mechanism for displaced responsibility. The concept is named and applied correctly (traits predict aggregated behavior, situations decide single occasions), and the response explains why a null result here counts as evidence rather than as missing data.

EBQ 3 · 45 minutes · 7 points

Using the sources provided, develop and justify an argument about whether beliefs about ability and control can be changed in ways that improve outcomes.

Directions

You have 45 minutes. Use the three sources below. Your evidence must come from at least two of them. Cite the sources you use as "Source A," "Source B," or "Source C," or by the authors' names. Write in complete sentences and label each part.

  • A. Propose a specific and defensible claim, based in psychological science, that responds to the question.
  • B(i). Support your claim using one specific and relevant piece of evidence from one of the sources, and cite that source.
  • B(ii). Explain how the evidence in Part B(i) supports your claim, using a psychological perspective, theory, concept, or research finding you learned in this course.
  • C(i). Provide additional support for your claim using a second specific and relevant piece of evidence, drawn from a different source than the one used in Part B(i), and cite that source.
  • C(ii). Explain how the evidence in Part C(i) supports your claim, using a psychological perspective, theory, concept, or research finding that is different from the one used in Part B(ii).
Source A

Source A: research summary written for this course, based on Mueller and Dweck (1998).

Across six experiments, fifth graders recruited from several schools worked a set of nonverbal reasoning problems and were then told they had done well. The only thing that differed between groups was the sentence that followed. Children randomly assigned to the ability condition were told they must be smart at these problems; children in the effort condition were told they must have worked hard; a comparison group was praised for the score alone. Offered a choice of what to do next, most effort-praised children took a challenging set described as one they would learn from, while most ability-praised children took an easier set described as one they would do well on. Every child then worked a much harder set that almost everyone failed. Afterward the ability-praised children reported enjoying the work less and wanting to stop, were more likely to attribute the failure to not being smart enough, and scored lower on a final set of easy problems than they had on the first set. The effort-praised children attributed the failure to not having worked hard enough and improved on the final set. Several hundred children took part across the six experiments, which repeated the pattern with different tasks and wording. Limitation: the sessions were short, ran outside normal instruction, and measured performance minutes rather than months later.

Source B

Source B: research summary written for this course, based on Seligman and Maier (1967) and Maier and Seligman (2016).

Dogs were assigned to three conditions: shocks they could switch off themselves, the same shocks delivered on the same schedule with no way to stop them, and no shocks at all. The second group received exactly the pattern of shock that the first group's own behavior produced, so the single difference between those two groups was control. All the animals were later placed in an apparatus where escaping required only stepping over a low barrier. Animals from the controllable condition escaped promptly. Most animals from the uncontrollable condition did not try, remaining still and absorbing the shock. The researchers named the effect learned helplessness and, in 1978 with Abramson and Teasdale, recast it for humans as explanatory style: the person who attributes bad events to internal, stable, and global causes is the one who stops trying. In 2016 the original researchers revised the account on neuroscience evidence. Passivity, they argued, is the default response of brain circuits to prolonged aversive events and is not learned at all; what an animal learns when it can act is that it has control, and that learning inhibits the default. Limitation: small groups of animals, a procedure that would face far stricter animal-care review today, and a step from escape latency to human belief that is an inference rather than a finding.

Source C

Source C: research summary written for this course, based on Yeager and colleagues (2019).

A preregistered randomized experiment tested a brief online program in a nationally representative sample of United States public high schools. Sixty-five schools took part, and more than 12,000 ninth graders were randomly assigned, within their own schools, to the program or to a control activity of the same length. The program was two online sessions totalling about 25 minutes, delivered during the school day, which presented the idea that intellectual ability can grow with effort and better strategies and asked students to write about it in their own words. The outcome was grade point average in core ninth-grade courses, taken from school records rather than from self-report, and recorded by people who did not know which condition a student had been in. Among lower-achieving students, those who received the program earned about 0.10 grade points more than the control students, a small average difference that held across the sample, and more of them enrolled in advanced mathematics the following year. The effect appeared in schools where peer norms supported taking on challenging work and was near zero where those norms did not. Limitation: the average effect is small, it did not appear among higher-achieving students, and its dependence on the school environment means the program should not be expected to work everywhere.

Your response
What a reader looks for on this prompt
  • A. The claim (1 point). The sources support a "yes, with conditions" claim, and the conditions are what make it defensible: "beliefs about ability and control can be changed by brief interventions, and the change improves real outcomes, but the gains are small and appear only for students who held the unhelpful belief and only where the environment lets the new belief pay off." A skeptical claim also works, "the belief changes reliably in the laboratory but the payoff outside it is too small and too conditional to justify treating mindset programs as a fix for achievement gaps", because Source C supplies the effect size and the moderation. "Mindset matters" earns nothing.
  • B(i) and C(i): the evidence (1 point each). Cite a result, not a topic. Usable pieces: one sentence of praise changing which task children chose, and ability-praised children scoring lower on the final easy set than on the first while effort-praised children improved (Source A); animals with prior control escaping promptly while most animals without it did not try, or the 2016 revision that control rather than helplessness is what gets learned (Source B); the 0.10 grade-point gain for lower-achieving students across 65 schools and more than 12,000 students, the increased enrollment in advanced mathematics, or the near-zero effect where peer norms did not support challenge (Source C). Both pieces cannot come from one source.
  • B(ii). The first explanation (2 points). Attribution theory and explanatory style fit Source A exactly: praise for ability supplies an internal, stable, global cause for success, which the child then applies to failure as "I'm not smart enough," while praise for effort supplies an internal but unstable and controllable cause. Achievement goal orientation, performance goals versus mastery goals, is an equally good fit for the task choice. One point rather than two if the response says the children "got a fixed mindset" without naming the attributional mechanism or connecting it to the specific result.
  • C(ii). The second explanation (2 points), with a different concept. Self-efficacy inside Bandura's reciprocal determinism fits Source C: a changed belief about one's own capability changes behavior, such as enrolling in harder mathematics, which changes the environment the student then learns in, which feeds back on the belief. Reciprocal determinism also explains the moderation directly: the person factor alone cannot carry the outcome, so the effect disappears where peer norms do not cooperate. Effect size versus statistical significance is the other strong move here: 0.10 grade points is a real but small difference, and saying so is what makes the claim defensible rather than promotional.
  • Which pairs count as "different." Attributional style, which is about the causes a person assigns to an outcome, paired with self-efficacy and reciprocal determinism, which is about the loop between belief, behavior, and environment, satisfies the rule. So does attributional style paired with intrinsic motivation and the overjustification effect, or achievement goal orientation paired with reciprocal determinism. What fails: "growth mindset" and "fixed mindset," which are two ends of one construct; "learned helplessness" and "pessimistic explanatory style," which are the same theory in its 1967 and 1978 wordings; and "self-efficacy" and "self-confidence." Any of those caps the second explanation at 1 point.
Show a 7/7 response

A Beliefs about ability and control can be changed, and changing them improves real outcomes: but the gains are small, they show up in the students who held the unhelpful belief to begin with, and they depend on an environment that lets the new belief pay off.

B(i) Source A reports that a single sentence of praise was enough to move fifth graders' behavior: most children told they must be smart chose an easier follow-up task, and after a hard set that nearly everyone failed they blamed their ability, said they enjoyed the work less, and scored lower on a final easy set than they had scored on the first one, while the children told they must have worked hard chose the challenging task and improved on that final set.

B(ii) Attribution theory explains why one sentence did that much work. An explanatory style is the habit of assigning causes along three dimensions (internal or external, stable or unstable, global or specific), and the praise handed each group a different cause for the same success. "You must be smart" is internal, stable, and global, so when the hard set arrived those children had only one place to file the failure: a fixed quality of themselves that the evidence now said was low. Nothing can be done about a stable cause, so they disengaged, and their scores fell. "You must have worked hard" is internal but unstable and controllable, so failure meant the effort was insufficient, which is a repairable cause, and those children pushed on. The manipulated variable was the sentence, not the children, so the belief is what changed and the behavior followed it.

C(i) Source C shows the same kind of change outside a laboratory and also shows its limits: in a preregistered randomized trial across 65 nationally representative high schools and more than 12,000 ninth graders, two online sessions of about 25 minutes raised the core-course grade point average of lower-achieving students by about 0.10 points and raised the number who later enrolled in advanced mathematics, but the effect was near zero in schools where peer norms did not support taking on hard work.

C(ii) Reciprocal determinism, with self-efficacy as the personal factor, accounts for both halves of that result. Bandura's claim is that personal factors, behavior, and environment each act on the other two, so a belief only produces an outcome by changing behavior that the environment then rewards. Raising a student's belief that she can handle harder mathematics changes a choice, she enrolls, and the harder course supplies the instruction that raises the grade, which in turn confirms the belief. Where classmates treat challenge-seeking as foolish, the same belief change produces no new behavior, the loop never closes, and the effect is near zero, which is exactly the pattern Source C found. The 0.10-point gain should be read as a small effect that reached significance in a very large sample, not as a transformation.

Where the points are earned
  • A: Defensible claim (1): The claim takes a position and states three conditions the response then defends: small gains, concentrated in students who held the unhelpful belief, dependent on the environment. It is falsifiable: if Source C had found uniform effects across all schools and all achievement levels, the claim would be wrong.
  • B(i): Evidence from a source (1): Specific results from Source A, cited by source: the task chosen after ability versus effort praise, the attribution each group made after the hard set, and the fall below baseline on the final set for ability-praised children against improvement for effort-praised children.
  • B(ii): Explanation with a concept (2): Names attribution theory and explanatory style, states the internal-stable-global dimensions correctly, and applies them to the exact wording of each praise condition and the exact result that followed. It also makes the experimental point that the sentence, not the child, was manipulated, which is what licenses the causal reading.
  • C(i): Evidence from a different source (1): The second piece comes from Source C rather than Source A, and is specific on every count: 65 schools, more than 12,000 ninth graders, about 25 minutes of intervention, 0.10 grade points among lower-achieving students, higher advanced-mathematics enrollment, and a near-zero effect where peer norms did not support challenge.
  • C(ii): Explanation with a different concept (2): Reciprocal determinism and self-efficacy are a different family of explanation from attributional style: one is a three-way causal loop among person, behavior, and environment, the other a scheme for the causes a person assigns to an outcome. The concept is named, defined, and applied to both the positive result and the null result, and the closing sentence reads 0.10 grade points as a small effect in a large sample, which keeps the claim's first condition honest.

EBQ 4 · 45 minutes · 7 points

Using the three sources, develop and justify an argument about how much confidence a court should place in a confident eyewitness's memory.

Directions

You have 45 minutes. Read the three sources, then respond to every part in complete sentences, in order, labelling each part. Your evidence must come from at least two of the three sources. Cite a source as Source A, Source B, or Source C, or by the researcher's name. You do not need outside research, but everything you say a source reports must be accurate.

  • A. Propose a specific and defensible claim based in psychological science that responds to the question.
  • B(i). Support your claim using at least one piece of specific and relevant evidence from one of the sources.
  • B(ii). Explain how the evidence from Part B(i) supports your claim using a psychological perspective, theory, concept, or research finding learned in AP Psychology.
  • C(i). Provide additional support for your claim using at least one additional piece of specific and relevant evidence from a different source than the one used in Part B(i).
  • C(ii). Explain how the evidence from Part C(i) supports your claim using a different psychological perspective, theory, concept, or research finding than the one used in Part B(ii).
Source A

Source A: research summary written for this course, based on Sperling (1960).

Adult observers viewed a grid of twelve letters, arranged in three rows of four, flashed on a screen for a fraction of a second. In the whole-report condition they wrote down as many letters as they could recall from the whole grid. They typically produced about four or five, and several reported that more letters had been visible but had faded before they could be written down. In the partial-report condition, a high, medium, or low tone sounded immediately after the display disappeared, and the pitch indicated which single row to report. Observers could not know in advance which row would be called. Cued this way, they reported roughly three of the four letters from whichever row the tone named. Because performance on a randomly chosen row was that high, the researcher inferred that nearly the whole array had been available for a brief moment after the display ended. Delaying the tone weakened the advantage; at a delay of about one second, partial report was no better than whole report. The design was an experiment in which each observer served in both conditions; the instruction was the manipulated variable and the number of letters correctly reported was the measure. Only a small number of practiced observers took part, each completing many trials. The results are treated as evidence for a very brief visual store, now usually called iconic memory, that holds more than can be reported before it decays.

Source B

Source B: research summary written for this course, based on Simons & Chabris (1999).

Observers watched a short video of two teams passing basketballs, one team in white shirts and one in black, and were asked to keep a silent count of the passes made by one of the teams. In some conditions the task was made harder by requiring separate counts of bounce passes and aerial passes. Partway through the video, a person in a gorilla suit walked into the middle of the scene, turned toward the camera, and walked out the other side, on screen for several seconds. Afterward each observer reported the count and was then asked whether anything unusual had appeared. Across conditions, roughly half of the observers failed to report the unexpected event; in the most-cited version of the report about 46% missed it. Misses were more common when the counting task was harder, and more common when observers were attending to the white-shirted team than the black-shirted one. Observers who watched the same video with no counting task noticed the figure nearly every time. Many of those who missed it were surprised on a second viewing and some questioned whether the same video had been shown. The design was an experiment: video version and counting difficulty were manipulated across groups, and the measure was simply whether the observer reported the unexpected event. Participants were college students tested individually in a laboratory.

Source C

Source C: research summary written for this course, based on Loftus & Palmer (1974).

Forty-five college students watched short films of traffic accidents and then completed a questionnaire about what they had seen. Buried in the questionnaire was one critical item asking how fast the cars had been going when they met; the verb in that sentence differed across groups and included "smashed into," "collided," "bumped," "hit," and "contacted." Every group had watched the same films. Average speed estimates tracked the verb: the group questioned with "smashed into" gave a mean of about 40.8 miles per hour, while the group questioned with "hit" gave a mean of about 34.0. A second experiment with new participants showed a brief film of a crash and then asked the speed question with "smashed into," with "hit," or not at all. A week later everyone returned and answered a further questionnaire that included a question about whether they had seen broken glass. There was no broken glass in the film. About 32% of those originally questioned with "smashed into" reported seeing it, compared with about 14% of those questioned with "hit." Both studies were experiments in which the wording of the question was the manipulated variable; the measures were a numerical speed estimate and, a week later, a yes-or-no report. The participants were undergraduates watching films, not witnesses to real crashes.

Your response
What a reader looks for on this prompt
  • A: A claim that takes a position on how much weight confident eyewitness testimony deserves: for example, that it should be treated as a lead requiring corroboration rather than as proof. "Memory is complicated" or a restatement of the question is too general to defend and earns nothing.
  • B(i): One particular finding, with its source named: the roughly 46% of observers in Source B who missed the unexpected figure, the three-of-four partial report in Source A, or the 40.8 versus 34.0 mile-per-hour estimates in Source C. Saying only that "Source B is about attention" is not evidence.
  • B(ii): A named concept, correctly applied, that connects that finding to the claim: selective attention and encoding failure, the limited capacity of short-term memory, or the rapid decay of sensory memory. The explanation must say why the finding makes confident testimony less trustworthy, not merely restate the finding.
  • C(i): A second specific finding drawn from a different source than B(i). Two findings from the same summary cannot earn this point, however good they are.
  • C(ii): A second named concept from a genuinely different family of explanation: reconstructive memory, the misinformation effect, source amnesia, or schema-driven retrieval if B(ii) used attention and encoding. "Attention" and "noticing" count as the same concept twice and cap this row at 1 point.
Show a 7/7 response

A A court should treat a confident eyewitness memory as a lead worth investigating rather than as proof, because memory can fail at two separate stages (what gets into the system, and what comes back out of it), and neither failure feels like a failure to the witness. Source A shows how briefly the raw visual record is even available. Confidence tracks how vivid a report feels, not how accurate it is, so testimony should carry real weight only when something outside the witness's memory corroborates it.

B(i) Source B reports that when observers were counting basketball passes, about 46% of them failed to notice a person in a gorilla suit who walked through the middle of the scene and was visible for several seconds, while observers given no counting task noticed the figure nearly every time. B(ii) This supports my claim through selective attention and the encoding failure it produces. Attention is a bottleneck: what is in the attended stream gets encoded into short-term memory, and what falls outside it is never stored, so there is no trace to retrieve. The observers were not careless; they were doing exactly what they were told, and the price of that focus was blindness to everything else. A witness watching a weapon may never encode the rest of the scene, including the face a lineup will later ask about. Because she has no memory of failing to see it, nothing signals that anything is missing, and she will describe what she did encode with complete sincerity. The jury hears certainty; the psychology says that certainty is about the fragment that was attended, not about the event.

C(i) Source C reports that a week after watching a filmed crash, about 32% of participants asked how fast the cars were going when they "smashed into" each other said they had seen broken glass, against about 14% asked with the verb "hit", and the film contained none. C(ii) This supports the same claim through a different mechanism: reconstructive memory and the misinformation effect. Retrieval is not playback. A memory is rebuilt each time from a partial trace plus whatever else fits it: schemas, expectations, and information met after the event. One verb implied a violent collision, broken glass fits that picture, and a week later the inference was reported as something seen. What should worry a court is that the distortion came from the question itself, asked by an experimenter with no motive to mislead. Every interview, lineup instruction, and news story between the crime and the trial is post-event information of the same kind, so a confident memory at trial may be partly a record of those rather than of the crime. Because Source B's failure happens at encoding and Source C's at retrieval, careful interviewing cannot repair what was never encoded, and close attention cannot protect a memory from later contamination, which is why corroboration, not confidence, should decide the weight.

Where the points are earned
  • A: Defensible claim (1): The first sentence takes a position the sources can defend and that someone could disagree with: confident eyewitness memory is a lead requiring corroboration, not proof, because memory fails at two stages. It answers the question asked ("how much confidence"), rather than summarizing the sources or asserting something no evidence could contradict.
  • B(i): Evidence from a source (1): A specific, accurate finding with the source cited: about 46% of observers in Source B missed the unexpected figure while observers without the counting task saw it nearly every time. The contrast between the two conditions is what makes it evidence rather than an anecdote.
  • B(ii): Explanation with a concept (2): Selective attention is named and applied correctly (the attended stream is what reaches short-term memory, so unattended information is never encoded and cannot be retrieved), and the explanation then carries that mechanism to the claim by pointing out that the witness has no internal signal that anything is missing, which is exactly why confidence and accuracy come apart.
  • C(i): Evidence from a different source (1): The second finding comes from Source C, not Source B: about 32% versus about 14% reporting broken glass that was never in the film, a week after a single verb changed. The evidence therefore spans two sources, as the directions require.
  • C(ii): Explanation with a different concept (2): Reconstructive memory and the misinformation effect are named and correctly applied to a retrieval-stage distortion, which is a genuinely different family of explanation from the attention and encoding account used in B(ii): different stage, different mechanism, not a rewording. The closing sentences make the difference explicit and use it to justify the claim, since a fix for one stage does nothing for the other.

EBQ 5 · 45 minutes · 7 points

Using the three sources, develop and justify an argument about whether self-control is best understood as a limited resource that can be used up.

Directions

You have 45 minutes. Read the three sources, then respond to every part in complete sentences, in order, labelling each part. Your evidence must come from at least two of the three sources. Cite a source as Source A, Source B, or Source C, or by the researcher's name. You do not need outside research, but everything you say a source reports must be accurate.

  • A. Propose a specific and defensible claim based in psychological science that responds to the question.
  • B(i). Support your claim using at least one piece of specific and relevant evidence from one of the sources.
  • B(ii). Explain how the evidence from Part B(i) supports your claim using a psychological perspective, theory, concept, or research finding learned in AP Psychology.
  • C(i). Provide additional support for your claim using at least one additional piece of specific and relevant evidence from a different source than the one used in Part B(i).
  • C(ii). Explain how the evidence from Part C(i) supports your claim using a different psychological perspective, theory, concept, or research finding than the one used in Part B(ii).
Source A

Source A: research summary written for this course, based on Baumeister and colleagues (1998).

Undergraduate participants were asked to skip a meal before the session so that they would arrive hungry. Each was seated at a table holding freshly baked chocolate-chip cookies and a bowl of radishes. Random assignment determined the condition: one group was told to eat only the radishes and to leave the cookies alone, a second group was permitted the cookies, and a third was brought to a table with no food on it. The experimenter then left the room, so resisting the cookies had to be done unobserved. Everyone next worked on a set of geometric tracing puzzles that had in fact been constructed so that they could not be solved, and the measure was how many minutes the participant kept working before giving up. Those who had been made to resist the cookies quit after roughly 8 minutes on average, while those who had been allowed the cookies kept working for about 19 minutes and those who had seen no food for about as long. The study was an experiment with random assignment, and persistence time on the unsolvable task served as the operational definition of self-control still available. The sample was a few dozen students, so the group difference, though large, rested on a small number of people. The researchers proposed that acts of self-control draw on one limited resource, a state later named ego depletion.

Source B

Source B: research summary written for this course, based on Hagger and colleagues (2016).

Nearly twenty years after the original report, a coordinated project set out to test the depletion effect with one procedure agreed in advance. Twenty-three laboratories in several countries ran the identical protocol, and the hypothesis and analysis plan were registered publicly before any data were collected, so that no laboratory could select an analysis after seeing its results. More than 2,000 participants in total were randomly assigned either to a letter-crossing task governed by a demanding rule, designed to require self-control, or to a version of the same task with no such rule; everyone then completed a second task on which self-control could be measured. Pooled across all the laboratories, the difference between the depleting and non-depleting conditions was about d = 0.04: an effect small enough to be indistinguishable from no effect at all. Individual laboratories varied, some reporting small differences in one direction and some in the other, the pattern expected from chance variation around zero. The report also observed that published studies of depletion had nearly all reported positive results, and that when journals accept positive findings more readily than null ones, a research literature can appear far stronger than the evidence behind it. The design was a set of direct replications rather than a new experiment, and whether the agreed procedure was the fairest test of the original idea remains disputed.

Source C

Source C: research summary written for this course, based on Mischel and colleagues, with the replication by Watts, Duncan & Quan (2018).

In the original work, a preschool child was seated alone at a table with a single treat (a marshmallow, a cookie, or a pretzel), and told that the experimenter was leaving for a while. If the child waited until the experimenter came back, she would receive two treats; if she did not want to wait, she could ring a bell at any time and eat the one in front of her. The measure was how many minutes the child waited. Waiting times varied a great deal. Follow-ups years later, of a few dozen children who had all attended one university's nursery school, reported that children who had waited longer tended to be rated as more competent adolescents and to score higher on college-entrance tests. A later study repeated the delay task with more than 900 children from a much broader and more socioeconomically varied sample and followed them to about age 15. The association between waiting time and later achievement was present but roughly half the size reported earlier, and it shrank further, in some analyses to near zero, once family background, mother's education, and early cognitive ability were statistically controlled. At the follow-up stage both are correlational: waiting time was measured rather than assigned, so the link to later outcomes cannot show that waiting caused them.

Your response
What a reader looks for on this prompt
  • A: A claim that takes a side on whether self-control behaves like a resource that is spent: for example, that the depletion account has not survived rigorous retesting and self-control is better explained by situation and learning history. A response arguing the other way can also earn every point, as long as it handles Source B honestly. "Self-control is important" or a restatement of the question is not defensible.
  • B(i): One particular finding, with the source named: the near-zero pooled effect of about d = 0.04 across 23 laboratories and more than 2,000 participants in Source B, the roughly 8 versus 19 minutes of persistence in Source A, or the halved association in Source C. "Source B failed to replicate" without a figure or a described result is thin.
  • B(ii): A named concept, correctly applied: effect size, direct replication, preregistration, publication bias and the file drawer, or the ego-depletion strength model itself. Translating d = 0.04 into what it means about two distributions is the move that earns the second point.
  • C(i): A second specific finding from a different source than B(i). Two figures out of the same summary cannot earn this point.
  • C(ii): A concept from a different family than B(ii). If B(ii) argued from methodology, C(ii) needs a substantive psychological account: operant conditioning and reinforcement history, delay of gratification, executive function, or the person-situation controversy. Using "replication" for one part and "effect size" for the other is the same concept twice and caps this row at 1 point.
Show a 7/7 response

A Self-control is not well described as a single limited resource that gets used up. The better-tested evidence points to self-control as behavior that depends on the situation and on what a person's environment has taught them to expect about waiting. Source A is the reason the resource idea took hold, and it is a real experiment, but a claim about human nature has to rest on what replicates rather than on what was found first.

B(i) Source B reports that 23 laboratories ran one agreed, preregistered procedure with more than 2,000 participants and found the difference between the depleting and non-depleting conditions to be about d = 0.04, with individual laboratories scattered in both directions. B(ii) This supports my claim through the research-methods concepts of effect size and direct replication. An effect size of 0.04 means the two groups' means sit about four-hundredths of a standard deviation apart, so the distributions lie almost on top of each other and there is nothing a real person would notice. Sample size makes that estimate trustworthy: with more than 2,000 participants it is precise, so an effect anywhere near the size of the gap in Source A would have been hard to miss. Laboratories landing on both sides of zero is what chance variation around no effect looks like. Because the analysis plan was registered in advance, no one could keep testing until something worked. Publication bias explains the gap between the two sources: journals print positive results far more readily than null ones, so a literature built from small studies can look overwhelming while the real effect is near zero.

C(i) Source C reports that when the delay-of-gratification task was repeated with more than 900 children from a far more socioeconomically varied sample, the association between waiting time and later achievement was about half the size found in the original follow-ups, and shrank further, in some analyses to near zero, once family background, mother's education, and early cognitive ability were controlled. C(ii) This supports my claim through a different kind of explanation: the behavioral perspective, and specifically reinforcement history. Waiting for the second treat is operant behavior, and it only pays if the promise is reliable. A child whose household has consistently delivered what it promised has been reinforced for waiting; a child whose environment makes promises uncertain is behaving sensibly by taking the sure thing now. That is why controlling for family background absorbs most of the association: waiting time measures the child's learning history and circumstances as much as anything stored inside the child. The person-situation controversy makes the same point from the trait side: traits predict behavior averaged over many occasions, not one act, and a single afternoon with a single marshmallow is one act. If the score largely indexes the situation, the inner fuel tank explains very little, and the way to improve self-control is to change what the situation rewards rather than ration a supply.

Where the points are earned
  • A: Defensible claim (1): The response takes a position the sources can be used to defend and that someone could argue against: self-control is not a limited resource that is spent, but behavior shaped by situation and learning history. It answers the question asked, and it stays defensible by conceding that Source A is a real experiment rather than pretending the opposing evidence does not exist.
  • B(i): Evidence from a source (1): A specific, accurate finding with the source cited: the pooled effect of about d = 0.04 from 23 laboratories and more than 2,000 participants in Source B, together with the scatter of individual laboratories in both directions.
  • B(ii): Explanation with a concept (2): Effect size and direct replication are named and applied correctly. The response translates d = 0.04 into overlapping distributions, explains why a large sample makes the near-zero estimate precise enough to rule out an effect the size of Source A's, and adds preregistration and publication bias to explain why the published record looked stronger than the evidence. That is a mechanism connecting the evidence to the claim, not a restatement of it.
  • C(i): Evidence from a different source (1): The second piece of evidence comes from Source C rather than Source B: the association between waiting time and later achievement was about half the size in a sample of more than 900 children and fell further, in some analyses to near zero, once family background and early cognitive ability were controlled. The evidence therefore spans two sources.
  • C(ii): Explanation with a different concept (2): The behavioral perspective, waiting as operant behavior under the control of reinforcement history and the reliability of the promised reward, is a genuinely different family of explanation from the methodological concepts used in B(ii), and it is applied to the specific result that the association shrinks when background is controlled. The person-situation controversy is then used correctly: traits predict aggregated behavior, not one act on one afternoon. Together they explain why the delay score reflects the child's circumstances rather than a fixed internal store, which is exactly the claim.

Unit recap

Unit recap

0:00 / 0:00

Animated recap with on-screen narration. Turn on Voice to have it read aloud (uses your device's built-in voice). Pressing play counts as your one free video.

Free preview complete

That's the end of the free preview.

You've opened five lessons (or watched a unit video), which is as much as we can show without a subscription. Everything you've already opened stays available; use the outline on the left to go back to it.

Book a tutor instead