Hawkeye said:
In the context of this ongoing discussion the analogy is totally relevant and speaks directly to one of the issues of this conversation. "Is A-B testing as reliable and conclusive as many in this thread assume it is?" It appears that the conclusion among many who make professional sound their business is that it is only one measurement that has limitations and certainly not the final measurement.
Manley labs is well regarded in both pro and audiophile circles as a manufacturer of high-end, very good sounding equipment for recording and for home. It's full of engineers who, as paulraphael has pointed out, have discovered that their are limitations to the type of test that everybody seems to be holding up so high as a final arbiter of comparison. They rely also on extended listening to familiar tracks or passages to let the ear be the final measurement.
A secondhand loose recollection of what engineers at
one manufacturer said is "many"? How do you get to that conclusion?
And it's by no means clear that these engineers have "discovered" anything of the sort. They believe it, perhaps, or assert it, but discover? Not unless they can prove it.
The problem with holding up the extended listening approach as a superior alternative to ABX testing--and this is really basic to any kind of study that attempts to establish a matter of fact--is that it doesn't control for confounding variables. In the first place, when you know what you're listening to, it is impossible to guarantee that you're not subject to assumptions and predispositions. (This doesn't guarantee that you will be prejudiced either; it just can't rule out the possibility, which you need to do to ensure the validity of your results.) In the second place, we know that people lose their ability to recall small differences in sound over time. Therefore, the longer you have preamp B in your house, listening exclusively to it, the less you'll be able to recall any minute differences it may have had with preamp A. (It's of course quite possible to remember gross, obvious differences. If preamp A had 20% distortion, and preamp B had 0.02%, you'd have no problem remembering that.)
Finally, the whole idea of one piece of gear sounding better than another is that you can hear the differences. That's what sounding better means. Thus, if you can't hear the audible differences, there aren't any audible differences, by definition. I seriously don't see how anybody can claim that extended listening could reveal any differences that would not also be apparent on a sufficiently rigorous, sufficiently extensive ABX test. If there are differences you can hear with extended listening, differences substantial enough to convince you that one piece of gear is better than another, why would these audible differences disappear in the context of ABX testing? Shouldn't they still be audible if they're real?
Again, for the record, I certainly don't believe all pieces of gear sound alike. I am certain that in many cases blind testing will reveal quite obvious differences. I've heard it. I just don't buy the idea that the extended in-home testing many audiophiles like to advocate is a better tool for the job.
You also have to keep in mind that in addition to being in the business of making gear, audio companies are in the business of
selling it. That in no way makes them automatically liars or crooked--I would in fact bet that most people who go into high-end audio do so out of a true passion for it--but the fact that they have an interest in the situation is a potential confounding variable.
Finally (really, I promise), as for the ear being the final measurement, that's exactly what happens in an ABX test: your ear is the measuring instrument. You hear a difference or you don't. The thing is, in this setting, the things your ear is supposed to measure are presented impartially and equally, with as many variables as possible controlled for. In the scenario you advocate, the ear is also the measuring instrument, but the things measured are not presented equally and impartially. You are evaluating current experience with a piece of gear whose maker and model you know against an increasingly dim recollection of another piece of gear, whose maker and model you also know, in the past. It simply is not as good and reliable a test.