Friday, October 09, 2026

Can Shelter Dog Behavior Tests Predict Behavior After Adoption?

Shelter behavior evaluations can provide useful observations about what a dog did under conditions, but research does not support treating a one-time shelter test as a reliable forecast of how an individual dog will behave after adoption. A shelter behavior evaluation can tell us something important: what a dog did during that evaluation.

 

What it cannot automatically tell us is what the same dog will do weeks or months later in a home, with different people, different relationships, different resources, different opportunities, and a different learning history.

 

Research does show some shelter-to-home correspondence for certain behaviors. But the evidence is substantially weaker when we ask the harder question: Can a shelter test accurately predict what a particular dog will do after adoption? For several outcomes shelters care about most, including behaviors classified as aggression, resource guarding, and separation-related behavior, the answer is often no, or at least not with the degree of certainty that consequential decisions would require.

 

How strongly does shelter behavior correlate with behavior after adoption?

 

Some correspondence exists, but the size and usefulness of that relationship depend on exactly what was measured. Poulsen, Lisle, and Phillips assessed shelter dogs with a multi-part behavior evaluation and later repeated the assessment after adoption for 39 dogs. The second assessment occurred an average of about 81 days later. The correlation between the total shelter score and the later assessment was r = 0.29. The study also found little shelter-to-home correspondence for several tests involving direct physical interaction with an evaluator.

 

That does not prove that shelter and home behavior are unrelated. It tells us something narrower and more useful: in that sample, the overall shelter score showed limited correspondence with the later assessment. More importantly, a correlation coefficient is not the same thing as individual predictive accuracy. A study can find an association across a group of dogs while still misclassifying many individual dogs. Patronek, Bradley, and Arps later identified exactly this problem in the shelter-assessment literature. Correlation, regression, statistical significance, agreement, reliability, and individual predictive accuracy are different statistical questions.

 

What is the shelter trying to predict?

 

A behavior test cannot simply be described as “predictive” without identifying what outcome is being predicted. Predicting whether a dog will ever growl is different from predicting whether a dog will bite. Predicting a bite is different from predicting an injurious bite. Predicting one incident is different from predicting repeated attacks.

 

Predicting whether an adopter will report resource guarding is different from predicting whether the dog can safely live in a particular household. The time period matters too. Are we predicting behavior during the first week after adoption? Thirty days? Six months? Several years?

 

A prediction only makes sense when we specify the behavior, circumstances, and time horizon. That sounds obvious, but much of the confusion surrounding shelter evaluations begins when different outcomes are collapsed into a broad word such as “aggression” or “temperament.”

 

Does a shelter observation tell us something real about the dog?

 

Yes. An observation should not be discarded merely because its predictive value is limited. If a dog freezes when a person reaches toward its food bowl, that happened. If a dog avoids an unfamiliar person during handling, that happened. If a dog readily approaches people during several social interactions, that happened. If a dog growls during restraint, that happened. 

 

The scientific problem begins with the next translation. A dog growled during a particular food test can become: The dog is a resource guarder. That can become: The dog will guard food in a home. And that can become: The dog is dangerous. Those are not four ways of saying the same thing. They are four different claims, and each requires additional evidence. The first is an observation. The second is a classification or inferred trait. The third is a prediction. The fourth is a safety judgment.

 

Keeping those operations separate is one of the most important things we can do when evaluating shelter dogs.

 

Are some behaviors more predictable than others?

 

Some studies have found greater shelter-to-home correspondence for friendly, social, fearful, or anxious behavior than for several categories of problem behavior.

 

Mornement and colleagues evaluated the B.A.R.K. protocol and found that it predicted some fearful and friendly behavior after adoption. It did not successfully predict post-adoption problem behavior or aggression. Clay and colleagues later studied 123 adopted shelter dogs. Friendly or social behavior, fear, and anxiousness measured in the shelter predicted corresponding post-adoption measures, while aggression, food guarding, and separation-related behaviors were not reliably predicted.

 

We should not turn those findings into another universal rule. They do not establish that broad traits are always stable or that specific behaviors are always unpredictable.  They show that some studies have found better correspondence for certain behavioral dimensions than for several of the problem behaviors shelters most want to forecast.

 

How well does a shelter test predict resource guarding?

 

Resource guarding is one of the clearest examples of why a group-level association is not the same thing as accurate prediction for an individual dog. Marder and colleagues followed 97 adopted dogs. Twenty had been classified as showing food-related aggression during the shelter evaluation. Of those 20 dogs, 11 were later reported by adopters to show food-related aggression. Nine were not. Among the 77 dogs classified as negative during the shelter assessment, 17 were subsequently reported by adopters to show food-related aggression in their homes. The shelter result and adopter reports were statistically associated. But the positive shelter result clearly did not produce certainty about what adopters would later report.

 

A 2020 study by Betty McGuire, Destiny Orantes, Stephanie Xue, and Stephen Parry produced a similar pattern. Shelter resource-guarding assessments were significantly associated with later adopter reports for several situations, but the positive predictive values were low. More than half of the dogs identified as resource guarding by either the shelter evaluation or surrendering owner did not show corresponding guarding in the adopter reports. Agreement among surrender information, shelter evaluation, and adopter reports was only in the fair range when all three were available. That does not mean the shelter observations were false. A dog may genuinely guard a resource during a shelter test and not guard a resource under the circumstances later encountered in its home. That distinction is critical.

 

Does “not reported after adoption” mean the behavior was absent?

 

Not necessarily. Most post-adoption validation studies do not install researchers in the adopter's home and continuously observe the dog. They rely heavily on adopter questionnaires, interviews, or rating instruments. Those are legitimate sources of evidence, but they are still measurements. An adopter may never try to remove the dog's food. A dog adopted into a home without children may never encounter the circumstances necessary to observe child-directed behavior. A dog with limited contact with other dogs may never have an opportunity to display dog-directed behavior. A dog may possess bones in one home and never receive bones in another.

 

Therefore: No behavior reported does not always mean the dog has demonstrated that it would not perform the behavior. Sometimes the correct conclusion is simply:

 

We do not know because the relevant situation was not observed.

 

This also means that a shelter-positive, home-negative result should not automatically be called proof that the shelter observation was a false positive. The two environments may simply have presented different conditions.

 

Could the shelter's response to a test result change the later outcome?

 

Yes, and this makes predictive research even more complicated. Imagine that a dog displays resource guarding during a shelter assessment. The shelter then tells the adopter about the behavior, chooses a home without children, gives the adopter management instructions, recommends avoiding unnecessary removal of high-value resources, provides behavior support, or performs behavior modification before placement. If the adopter later reports no guarding problem, what does that mean? One possibility is that the shelter test overpredicted the problem. Another is that the behavior was real but the later environment, handling, management, learning, or intervention reduced the opportunity for it to occur. A third possibility is that guarding occurred but was not recognized or reported. Those explanations cannot always be separated after the fact.

 

This is why predictive validity is harder to establish than simply comparing a shelter score with a later questionnaire.

 

What does “aggression” mean in shelter research?

 

It depends on the study, and this is one of the biggest sources of misunderstanding. Researchers often combine several behaviors within a category called aggression. Those behaviors can include freezing, stiffening, showing teeth, growling, barking, lunging, snapping, and biting. Those are not identical events. A growl is not an injurious bite. A lunge ending without contact is not the same outcome as an attack involving repeated bites. A snap is not automatically evidence of serious violence.

 

Christensen and colleagues followed 67 adopted dogs that had passed a shelter temperament test. After adoption, 40.9% were reported to have lunged, growled, snapped, and/or bitten. That finding demonstrated that the shelter test failed to anticipate a substantial amount of behavior included within the study's aggression category. It does not establish that 40.9% of those dogs later bit someone, caused injury, or committed a dangerous attack.

 

Whenever a study uses the word aggression, the finding should be interpreted according to that study's operational definition. It should not automatically be translated into violence, attack, dangerousness, biting, or injury.

 

What happens when dogs that fail an evaluation are never adopted?

 

Then we cannot know what those dogs would have done in adoptive homes.

 

Bollen and Horowitz analyzed behavioral evaluations from 2,017 shelter dogs. Dogs that failed the evaluation were not placed for adoption. The researchers therefore could not prospectively determine whether those dogs actually would have shown the predicted aggression after adoption. This creates an important research problem. Suppose a test identifies a group of dogs as high risk. Those dogs are withheld from adoption or euthanized. They will never produce post-adoption data. We cannot count their absence from later bite reports as proof that the test worked. But we also cannot assume they would all have been false positives. Their post-adoption behavior is unknown because adoption never occurred. The assessment helped determine whether the outcome could ever be observed. That makes clean validation especially difficult.

 

How large is the false-positive problem?

 

Published research suggests that it can be substantial, but even these numbers require careful interpretation.

 

Patronek, Bradley, and Arps reviewed more than 25 years of shelter behavior-evaluation research. They concluded that no canine behavior evaluation or individual subtest had approached accepted standards sufficient to justify calling it validated for routine shelter use. Across the study populations they reviewed, the mean reported false-positive error rate was 35.1%. When they estimated performance under prevalence conditions they considered more typical of shelter populations, the estimated false-positive rate was 63.8%. That second number deserves a qualifier. The 63.8% figure was a modeled estimate, not a directly observed national false-positive rate among shelter dogs.

 

The larger lesson is more important than either percentage. A test can perform reasonably in a selected research sample yet produce many incorrect positive classifications when used in a population where the target outcome is uncommon.

 

This is one reason Patronek and Bradley argued that predicting relatively uncommon, serious future behavior is extraordinarily difficult even when a test appears to have respectable statistical properties.

 

Why does the base rate matter?

 

Because uncommon outcomes are difficult to predict without generating large numbers of false alarms.

 

Suppose a test is trying to predict a serious injurious attack. If such attacks are uncommon among the population being tested, most dogs tested will never produce that outcome. Even a test that correctly identifies many true cases can still identify many more dogs who never produce the predicted event simply because the noncase population is so much larger. This is a basic feature of screening and predictive testing, not something unique to dogs. It also explains why the outcome has to be defined precisely.

 

Predicting any growl is statistically different from predicting a medically significant bite. Predicting one resource-guarding incident is different from predicting a dog that cannot be safely placed in a normal household. The more serious and uncommon the target event becomes, the more carefully a positive prediction must be interpreted.

 

Does behavior stay the same after adoption?

 

Not necessarily.

 

A 2023 study by Bohland and colleagues followed 99 adopted dogs from five Ohio shelters. Owners completed the C-BARQ questionnaire at approximately 7, 30, 90, and 180 days after adoption. Several measured C-BARQ behavior scores changed over time. Stranger-directed aggression scores, excitability, touch sensitivity, training difficulty, and chasing increased at one or more later time points, while separation-related behavior and attachment or attention-seeking decreased. That study was not primarily a validation study of a pre-adoption temperament test. Its relevance here is different. It demonstrates that post-adoption behavioral measurements themselves can change with time. So even after a dog enters a home, there may not be one permanent “post-adoption behavior” against which a shelter test can be compared. The dog at seven days, the dog at 90 days, and the dog at 180 days have experienced different amounts of learning, familiarity, relationship development, environmental exposure, and reinforcement.

 

Does this mean dogs have no stable behavioral tendencies?

 

No. The question is not whether individual dogs differ from one another or whether some behavioral tendencies persist. The question is whether a particular observation under shelter conditions accurately predicts a particular later outcome under different conditions. Behavioral consistency should be demonstrated rather than assumed from one event. A dog may readily approach strangers yet dislike restraint. A dog may guard food but not toys. A dog may interact comfortably with dogs in one setting and display threatening behavior behind a barrier. A dog may tolerate handling by familiar people but resist handling by strangers.

 

A single label such as “friendly,” “aggressive,” “dominant,” or “good temperament” can conceal those differences.

 

Should shelters stop doing behavioral observations?

 

No. The research supports being more careful about what those observations are asked to prove. Shelter staff should observe dogs. Foster caregivers can provide valuable information. Previous owners may provide useful history. Veterinary handling can reveal information. Volunteers, trainers, playgroup staff, and adopters may all observe different parts of the dog's behavioral repertoire. Structured assessments can also add observations.

 

The mistake is not collecting information. The mistake is turning one source of information into an oracle.

 

Current ASPCA policy reflects that distinction. The organization recommends collecting behavioral information from multiple sources and states that behavior assessments have not proved highly accurate or precise for predicting aggression after adoption. It also recommends that, except in cases it regards as egregious, concerning behavior during an assessment be corroborated in another environment before being used as the sole basis for a euthanasia decision.

 

That is a policy recommendation informed by the evidence. It should not itself be confused with the empirical evidence.

 

What should an adopter ask instead of “Did the dog pass the temperament test?”

 

Ask what the dog did. What happened immediately before the behavior? Who was present? Was the dog eating, resting, being restrained, playing, meeting another dog, or being approached by a stranger? Was there a growl, freeze, snap, bite, or injury? Did the behavior occur once or repeatedly? Has it been observed by more than one person? Did it occur only during formal testing or also during normal daily interactions? Was the same behavior seen in foster care? What happened after the event? Did the dog move away, recover, escalate, or repeat the behavior? What situations have never been observed?

 

Those questions preserve information that labels throw away. “Failed aggression test” tells me far less than a careful description of what the dog actually did.

 

When does professional assessment make sense?

 

Professional assessment makes sense when the available information could affect safety but different sources conflict, important details are missing, or a shelter label has been treated as though it explains more than the underlying observations establish.

 

A dog that gave one growl during a provocative food test presents a different assessment problem from a dog with several independently documented injurious bites. A dog that avoids handling in a kennel but readily interacts with people in a foster home gives us important evidence that context changes the response.

 

An experienced professional can compare the conditions under which behavior succeeds and fails, distinguish observations from hypotheses, identify missing information, and help determine what conclusions the available evidence actually supports. The purpose is not to manufacture certainty where certainty is impossible. It is to reduce uncertainty enough to make a better decision.

 

If you have adopted a dog and the behavior you are seeing at home does not match what was reported in the shelter, an in-home assessment can be useful because the relevant behavior can be evaluated in the environment where it is actually occurring.

 

You can learn more about working with me at SamTheDogTrainer.com.

 

What is the most important distinction?

 

A shelter behavior evaluation is evidence about what a dog did under the conditions in which the dog was evaluated. That evidence may be useful. It may sometimes correspond with later behavior. It may identify something worth investigating further. But it should not automatically become a diagnosis of temperament, a prediction of future behavior, or a judgment that a dog is safe or dangerous.

 

The scientifically defensible sequence is: Observe the behavior. Preserve the context. Gather additional evidence. Define exactly what future outcome is being considered. Then decide how much confidence the evidence deserves.

 

A dog should not have to become a prediction simply because somebody performed a test.

 

Glossary

 

Behavior assessment: A structured procedure in which a dog's responses to specified situations or stimuli are observed and recorded.

 

Behavioral history: Information about previous behavior obtained from owners, finders, foster caregivers, shelter personnel, volunteers, veterinary professionals, or other observers.

 

Base rate: How common an outcome is in the population being evaluated. Base rates strongly affect the probability that a positive prediction is correct.

 

Correlation: A statistical measure of the degree to which two variables vary together. A correlation does not by itself demonstrate accurate classification or prediction for an individual dog.

 

False negative: A negative test result when the defined target outcome later occurs.

 

False positive: A positive test result when the defined target outcome is actually absent. This term requires caution when later behavior was not directly observed or the relevant situation never occurred.

 

Negative predictive value: The proportion of negative test results for which the defined outcome is actually absent.

 

Positive predictive value: The proportion of positive test results for which the defined outcome actually occurs.

 

Predictive validity: The degree to which a measurement predicts a specified later outcome.

 

Reliability: The consistency or repeatability of a measurement. A measurement can be reliable without accurately predicting the outcome of interest.

 

Resource guarding: Behavior associated with retaining access to a valued resource. Research definitions vary and may include postural changes, vocalizations, lunging, snapping, or biting.

 

Sensitivity: The proportion of actual cases correctly identified by a test.

 

Specificity: The proportion of actual noncases correctly identified by a test.

 

Temperament: A broad term referring to relatively enduring behavioral tendencies. A temperament label should not be assumed to predict identical behavior across situations.

 

Bibliography

 

Bohland, K. R., Lilly, M. L., Herron, M. E., Arruda, A. G., & O'Quin, J. M. (2023). Shelter dog behavior after adoption: Using the C-BARQ to track dog behavior changes through the first six months after adoption. PLOS ONE, 18(8), e0289356. doi:10.1371/journal.pone.0289356.

 

Bollen, K. S., & Horowitz, J. (2008). Behavioral evaluation and demographic information in the assessment of aggressiveness in shelter dogs. Applied Animal Behaviour Science, 112(1-2), 120-135. doi:10.1016/j.applanim.2007.07.007.

 

Christensen, E., Scarlett, J., Campagna, M., & Houpt, K. A. (2007). Aggressive behavior in adopted dogs that passed a temperament test. Applied Animal Behaviour Science, 106(1-3), 85-95. doi:10.1016/j.applanim.2006.07.002.

 

Clay, L., Paterson, M. B. A., Bennett, P., Perry, G., & Phillips, C. C. J. (2020). Do behaviour assessments in a shelter predict the behaviour of dogs post-adoption? Animals, 10(7), 1225. doi:10.3390/ani10071225.

 

Marder, A. R., Shabelansky, A., Patronek, G. J., Dowling-Guyer, S., & D'Arpino, S. S. (2013). Food-related aggression in shelter dogs: A comparison of behavior identified by a behavior evaluation in the shelter and owner reports after adoption. Applied Animal Behaviour Science, 148(1-2), 150-156. doi:10.1016/j.applanim.2013.07.007.

 

McGuire, B., Orantes, D., Xue, S., & Parry, S. (2020). Abilities of canine shelter behavioral evaluations and owner surrender profiles to predict resource guarding in adoptive homes. Animals, 10(9), 1702. doi:10.3390/ani10091702.

 

Mornement, K. M., Coleman, G. J., Toukhsati, S. R., & Bennett, P. C. (2015). Evaluation of the predictive validity of the Behavioural Assessment for Re-homing K9's (B.A.R.K.) protocol and owner satisfaction with adopted dogs. Applied Animal Behaviour Science, 167, 35-42. doi:10.1016/j.applanim.2015.03.013.

 

Patronek, G. J., & Bradley, J. (2016). No better than flipping a coin: Reconsidering canine behavior evaluations in animal shelters. Journal of Veterinary Behavior, 15, 66-77. doi:10.1016/j.jveb.2016.08.001.

 

Patronek, G. J., Bradley, J., & Arps, E. (2019). What is the evidence for reliability and validity of behavior evaluations for shelter dogs? A prequel to “No better than flipping a coin.” Journal of Veterinary Behavior, 31, 43-58. doi:10.1016/j.jveb.2019.03.001.

 

Poulsen, A. H., Lisle, A. T., & Phillips, C. J. C. (2010). An evaluation of a behaviour assessment to determine the suitability of shelter dogs for rehoming. Veterinary Medicine International, 2010, 523781. doi:10.4061/2010/523781.

 

ASPCA. Position Statement on Shelter Dog Behavior Assessments. Current organizational policy accessed October 2026.

No comments: