Replication: The upside of down

Replication
Science
Decision Making
Author

MSM

Published

September 21, 2026

Original work

Barone, M. J., Coulter, K. S., & Li, X. (2020). The upside of down: Presenting a price in a low or high location influences how consumers evaluate it. Journal of Retailing, 96(3), 397-410.

Abstract

Can changing the vertical location of a price (e.g., presenting it above or below a product image in an advertisement or retail display) influence consumer response? Drawing from conceptual metaphor theory, we propose that a price’s vertical location can activate metaphors that relate vertical locations to magnitudinal concepts. These “down = less” and “up = more” metaphors can subsequently influence evaluations of a target price as being monetarily low or high in magnitude. Consistent with this premise, several lab and field investigations demonstrate that prices provided in low (vs. high) locations lead to lower price perceptions, more favorable purchase intentions, and higher in-store sales.

Team BA autumn 2025

  • Patrick Bargetze
  • Alina Bösiger
  • Sebastian Löpfe
  • Anastasiia Lykholai
  • Sarina Menning
  • Flavian Römer
  • Filip Stojanovic

The idea

Do we read a price the way we read a thermometer? Up means more, down means less. Barone and colleagues asked what happens when that metaphor is quietly attached to a price tag: put the same price below a product rather than above it, and the “down = less” association should make it feel like a smaller number. Nothing about the price changes – only where on the screen the number sits.

That is the entire manipulation. Figure 1 shows one of our stimuli, the identical product at the identical price of CHF 7.00, in the two conditions participants could see.

(a) Price above the product
(b) Price below the product
Figure 1: The price location manipulation. Same product, same price, different vertical position.

We replicated four of the original studies. Everyone saw all four, and was randomly assigned to a price location within each one. The three product studies were shown in a random order, with the metaphor questions always last:

  • Study 1a – no price is shown. Participants write down the price they think is appropriate, in an empty box placed either above or below the product image.
  • Study 1b – a price is shown. Participants rate how low or high they find it (1 = very low, 9 = very high).
  • Study 2 – a price is shown. Participants rate how likely they would be to buy the product (1 = very unlikely, 9 = very likely).
  • Study 5 – no products. Participants report directly how strongly they hold the “up = more” and “down = less” associations.

We crossed price location with product type in studies 1a, 1b and 2: five everyday products (cheese, toothpaste, instant coffee, a cleaning sponge, a power strip) and five luxury products (a Tesla, a Jura coffee machine, a cashmere jumper, Dior lip balm, a Roja diffuser). The original used a single product category per study, so the product factor is our addition – it lets us ask whether the effect survives across very different price ranges.

Because the three product studies were rotated, we can also check whether it matters when in the questionnaire a study appeared.

How we analyse this

Two features of the design decide the statistics, and both are easy to get wrong.

First, everyone answered five items per study, so the rows in our data are not independent – five of them belong to the same person. Treating them as independent would inflate the degrees of freedom roughly fivefold and shrink every p value with it. Every model below is therefore a linear mixed model with a random intercept per participant.

Second, product type has to be in the model rather than in front of it. Assignment was random but not stratified, so the two location conditions ended up with slightly different product mixes: in Study 1a, 53% of the estimates in the above condition were for luxury products against 45% in the below condition. Luxury items cost roughly twenty times more, so a few percentage points of composition difference is easily enough to fake a location effect. Testing location adjusted for product type removes that. The alternative – entering location first in a sequential (Type I) ANOVA – credits it with the imbalance and produces significant results that vanish the moment you control for what was actually being priced.

We report each effect as an F test with partial eta squared and its 95% confidence interval, plus the estimated marginal mean difference between the two locations, averaged over product type.

Sample

We recruited through word of mouth among friends of the student group. 357 people opened the questionnaire. Of those, 284 answered at least one of the three attention checks correctly, and 219 of them also reached the end of the questionnaire. Those 219 participants make up the analysed sample.

The attention checks were the usual instructed-response items – a question about the colour of the sun that asks you to tick “blue”, for instance. Requiring only one of three to be correct is a lenient rule, and it is the rule the original analysis used, so we kept it here rather than tightening it after the fact. It is worth knowing that it does most of its work in combination with the completion filter: of the 284 who passed a check, 65 dropped out before the end.

Table 1: Participants by gender.
Gender n Mean age SD age
Female 138 35.4 15.3
Male 78 33.0 14.0
Non-binary 3 82.0 24.0

The sample is skewed toward female and young (Table 1). The three non-binary participants are too few to interpret on their own, and their mean age is pulled around by a single very old respondent.

Study 1a – estimating a price from scratch

Participants saw a product with an empty box either above or below it, and wrote in the price they considered appropriate. If “down = less” is doing any work, the estimates written into a low box should come out smaller.

Free-text prices are messy. Ours ranged from CHF 0 to CHF 4,500,000, and the standard deviation in the luxury condition (198,120) is an order of magnitude larger than the mean (Table 2). A handful of people valuing the Tesla at several million francs is enough to swamp everything else.

Table 2: Study 1a: raw price estimates in CHF.
Position Product n Mean SD
above Everyday product 265 27.1 130.6
above Luxury product 295 28324.9 265566.7
below Everyday product 293 16.4 32.6
below Luxury product 239 9643.6 24298.8

The original paper hit the same problem and log-transformed the estimates, so we do the same. On the log scale a franc difference counts the same whether it sits on a tube of toothpaste or on a car.

Figure 2: Study 1a: estimated price by price location and product type. Mean with 95% CI, log scale.
Table 3: Study 1a: geometric means of the price estimates.
Position Product n Mean (CHF)
above Everyday product 264 10.10
above Luxury product 294 230.15
below Everyday product 290 9.52
below Luxury product 236 200.31

Figure 2 and Table 3 show estimates that are a shade lower in the below condition for both product types, which is the direction the original predicts. The test does not back it up: F(1, 1081) = 0.53, p = .465, \(\eta^2_p\) = .000, 95% CI [.000, .007]. The estimated difference is -9.3%, 95% CI [-30.3%, +18.0%] – prices in a low box come out lower on average, but the interval comfortably contains zero and both plausible directions. Product type, by contrast, does what you would expect of it (F(1, 1081) = 533.99, p < .001, \(\eta^2_p\) = .331, 95% CI [.288, .372]); people know a Tesla costs more than toothpaste.

Study 1a does not replicate.

Study 1b – judging a price that is shown

Here the price was given and participants rated how low or high it felt. This is the most direct test of the original claim, and the one where we expected the clearest answer.

Figure 3: Study 1b: perceived price magnitude by price location and product type. Mean with 95% CI (1 = very low, 9 = very high).
Table 4: Study 1b: perceived price magnitude.
Position Product n Mean SD
above Everyday product 225 5.72 1.96
above Luxury product 265 7.25 1.98
below Everyday product 219 5.70 1.84
below Luxury product 285 7.38 2.09

Nothing, and this time not even a hint of a direction. The two locations land on top of each other for everyday products (5.72 versus 5.70) and, if anything, the low price is rated higher for luxury products (7.25 versus 7.38), which runs against the prediction (Figure 3, Table 4). The model agrees: F(1, 222.8) = 0.13, p = .721, \(\eta^2_p\) = .001, 95% CI [.000, .022], an estimated difference of +0.06, 95% CI [-0.29, +0.41] on the nine-point scale. The product manipulation meanwhile works exactly as expected (F(1, 224.2) = 79.48, p < .001, \(\eta^2_p\) = .262, 95% CI [.170, .352]) – so the measure is clearly sensitive, just not to price location.

The original reports a significant effect here (F(1, 203) = 4.36, p = .04) on a sample half the size of ours. We have power on our side and we do not find the effect.

Study 1b does not replicate.

Study 2 – purchase intention

Same setup as 1b, but participants rated how likely they would be to buy rather than how expensive it looked. The original argued that a price which feels smaller should translate into stronger intentions to buy.

Figure 4: Study 2: purchase intention by price location and product type. Mean with 95% CI (1 = very unlikely, 9 = very likely).
Table 5: Study 2: purchase intention.
Position Product n Mean SD
above Everyday product 255 4.87 2.89
above Luxury product 268 1.75 1.60
below Everyday product 279 5.10 2.85
below Luxury product 265 2.00 2.11

This is the closest we come. Participants were somewhat more willing to buy when the price sat below the product, for everyday products (4.87 versus 5.10) and luxury products (1.75 versus 2.00) alike (Figure 4, Table 5). The estimated difference, +0.24, 95% CI [-0.18, +0.65] on the nine-point scale, runs in the predicted direction and is the largest of the three studies, but the interval still includes zero and the test is not significant: F(1, 215.9) = 1.25, p = .266, \(\eta^2_p\) = .006, 95% CI [.000, .042].

That is worth being precise about. A consistent direction across two product categories is the kind of thing that makes you want to believe an effect is there, and it may well be – our interval is compatible with an effect of up to two thirds of a scale point. It is also compatible with nothing at all, and with a small effect in the opposite direction. This study does not settle the question.

The low purchase intentions for luxury products most likely goes back to our sample of students - being asked whether they would buy a Tesla; a mean of about 2 on a 9-point scale is a statement about budgets rather than about design.

Study 2 does not replicate.

Study 5 – do people hold the metaphor at all?

The whole account rests on people associating up with more and down with less. Study 5 measures that directly, with four statements rated from 1 to 9. The first two ask what “up” and “down” bring to mind (1 = less of something, 9 = more of something); the last two ask for agreement that a lower position means less and a higher position means more.

Because every participant answered all four statements, these are within-person comparisons, and the mixed model uses that pairing.

Figure 5: Study 5: strength of the vertical metaphor across the four statements. Mean with 95% CI.
Table 6: Study 5: metaphor strength.
Item n Mean SD
“Up” means more 203 7.21 1.76
“Down” means less 190 2.77 1.87
Higher = more 196 7.03 2.00
Lower = less 202 6.33 2.51

The metaphor is there, and unlike everything above it is not in doubt (Figure 5, Table 6). “Up” sits at 7.21 and “down” at 2.77 on the same scale, a gap of 4.44, 95% CI [4.08, 4.80] on the nine-point scale (F(1, 391) = 590.08, p < .001, \(\eta^2_p\) = .601, 95% CI [.546, .649]). Asked directly, people also agree that higher means more (7.03) more strongly than that lower means less (6.33), a smaller but reliable gap of 0.68, 95% CI [0.33, 1.04] (F(1, 198.1) = 14.43, p < .001, \(\eta^2_p\) = .068, 95% CI [.016, .145]).

So the central ingredient the theory needs is present in our sample, at full strength. What does not follow is the behaviour it is supposed to produce.

Does it matter when a study appears?

Studies 1a, 1b and 2 were rotated, so each one landed first, second or third for roughly a third of participants. That gives us a free check on something replications rarely get to look at: whether the position of a study inside the questionnaire changes what it finds.

Figure 6: Answers by the position a study occupied in the questionnaire. Mean with 95% CI, collapsed across product type. Study 1a is on the log scale.
Table 7: Task order added to each model.
Study Effect of order Location x order
1a – price estimation F(1, 1079) = 0.70, p = .401, η²p = .001 [.000, .007] F(1, 1079) = 0.82, p = .366, η²p = .001 [.000, .008]
1b – price magnitude F(1, 221.1) = 1.28, p = .259, η²p = .006 [.000, .041] F(1, 221.0) = 0.71, p = .400, η²p = .003 [.000, .034]
2 – purchase intention F(1, 213.2) = 8.50, p = .004, η²p = .038 [.004, .101] F(1, 213.5) = 0.52, p = .472, η²p = .002 [.000, .032]

Two things fall out of Figure 6 and Table 7.

The first is that order matters in Study 2. Purchase intentions were markedly lower when Study 2 came first than when it came later (F(1, 213.2) = 8.50, p = .004, \(\eta^2_p\) = .038, 95% CI [.004, .101]). Studies 1a and 1b show nothing of the sort (F(1, 1079) = 0.70, p = .401, \(\eta^2_p\) = .001, 95% CI [.000, .007] and F(1, 221.1) = 1.28, p = .259, \(\eta^2_p\) = .006, 95% CI [.000, .041]). The most likely reading is warm-up: participants who had already worked through a block of prices had a rough sense of what these products cost, and answered the buying question less conservatively than those who met the Tesla cold.

The second is that order does not interact with price location in any of the three studies (all p > .36). Whatever the rotation is doing to the overall level of responses, it is doing it equally to the above and below conditions. So the rotation is neither creating nor concealing a location effect – the nulls above are not an artefact of when people saw what.

Worth keeping in mind if you run something like this yourself. A rotation costs nothing and buys you the ability to tell “my effect” apart from “my questionnaire”.

Where that leaves us

Table 8: Original findings and our replication.
Study Original Replication Outcome
1a – price estimation F(2, 88) = 3.55, p = .03 F(1, 1081) = 0.53, p = .465, η²p = .000 [.000, .007] not replicated
1b – price magnitude F(1, 203) = 4.36, p = .04 F(1, 222.8) = 0.13, p = .721, η²p = .001 [.000, .022] not replicated
2 – purchase intention F(2, 86) = 4.14, p = .02 F(1, 215.9) = 1.25, p = .266, η²p = .006 [.000, .042] not replicated
5 – metaphor strength held by participants F(1, 391) = 590.08, p < .001, η²p = .601 [.546, .649] replicated

Table 8 is not the result we expected when we started. People in our sample hold the “up = more, down = less” metaphor about as strongly as it is possible to hold anything on a nine-point scale. It simply does not travel to their prices. Estimating a price from scratch, judging a price on the page, deciding whether to buy – in all three the confidence interval for the location effect straddles zero.

The direction is not random, for what that is worth. Two of the three point the way the original predicted, and Study 2’s difference of about a quarter of a scale point is the sort of effect the original was describing. With 219 participants we could not separate it from noise, and an effect that needs more than this to show itself is not the effect the original reported.

Caveats. Our design differs from the original: we ran all four studies on the same participants rather than using a fresh sample for each, we recruited a Swiss convenience sample rather than US students and mTurk workers (although you probably would not want to do that anyway today), and we added a product-type factor the original did not have. Any of those could explain why we see nothing where the original saw something. The rotation at least lets us rule out running order as the culprit.

Price location may still do something. On this evidence, whatever it does is smaller than a study of this size can see.

Data

The anonymised responses behind every figure and table on this page are in BAHS2025_public.csv: one row per answer, with the participant’s random ID, the study, the price position, the product type (price: A for everyday, L for luxury), the item (level), the serial position of that study in the questionnaire (order) and the answers themselves. Gender and age are held back and are not published.