Hi Tom,
Those are fair questions. The reason most casual comparative listening tests are flawed is that it's relatively trivial to set up two items next to each other and alternately try them out, but it's not trivial to set up their gains to match, so that the comparison is of performance-based sound quality, not loudness-based sound quality.
The most accurate way is to put a fixed signal level into both, measure their outputs, and adjust their gains so that the output voltages match. When I set up tests with an ABX comparator, I use its 1 volt 1 kHz sine wave oscillator to put a calibration signal into both amps, then I adjust each channel gain to put out the same voltage. I shoot for 20.0 volts, ±0.1 volt, so they are no more than 0.08 dB apart if one ends up at 20.1 and another at 19.9. Sometimes because of gain control detents I can't get exactly 20.0 volts on one or both amps, so I set the channels to a voltage that I can get on all of them, maybe 19.8 volts or 22.4 volts, whatever. The important thing is to get them as exactly the same as possible. This usually takes me maybe two to five minutes to fully calibrate two 2-channel amps for an ABX test.
You might not have a sine wave generator and an audio AC voltmeter to do this same sort of thing. Most people don't unless they are true audio geeks.

You might be able to get a decent approximation by using the gain control markings. Some QSC amps, like the PLX and PowerLight series, have markings that describe the actual voltage gain in dB. Other amps have markings labeled in dB below full gain, with "0" at full and maybe marks at -2, -4, -6, and so on. For something like this you need to know the amp's full gain. Some manufacturers standardize models or ranges of models on certain maximum gains; for example, the Crest Professional Series models had a standard maximum gain of 32 dB (40x); they might still have that, or maybe they might now have a switch for changing to another gain setting. Check the user documentation if you're not sure. Then you set the gain controls on the amps so that they match up; for example, -3 dB on an amp with a max gain of 32 dB, -7 dB on an amp whose max gain is 36 dB, and 29 dB on an amp that reads gain directly, are all the same thing: 29 dB. However, the precision of the gain controls and markings may be a little off from one amp to another, especially the further down you turn the controls. Imprecision will affect the objectivity of your comparison, but within reason it's still a lot better than a haphazard setup.
Ideally you would have some means of switching between one and the other so you don't have a long time interval between comparing each one to the other. Don't touch the amp gain controls, but instead turn up the source so that you can see how they compare when pushed; if B starts farting long before before A does, for example, it's good to know about so you can consider that in your evaluation. Turn down the sourece, too, so you can hear what low-level signals sound like; if A has noticeable hum or hiss or distortion that B doesn't, that's good to know about, too.
One thing about a double-blind test is that it's not for determining which one sounds better, but primarily for determining if they sound different. You have to first establish that they sound different before you can credibly state that one sounds better than the other. Maybe you can establish that they do sound different from each other, but one doesn't sound better than the other, or maybe you'll find that you can't distinguish one from the other by sound; that's not uncommon, especially for audio products that have essentially objective functions, like well-designed power amps.
A blindfold generally isn't necessary unless you want it.

An ABX comparator (
Link Removed) makes a double-blind test easy; the device randomly chooses X from A or B, and you can listen to X and A and B and compare as much as you want until you feel you know the identity of X, and you enter or record your decision. Then do it again with a new selection of X, and again and again until you have a series of tests that you can draw some statistically significant results from.
But even fewer people have ABX comparators than have audio generators and voltmeters. So without one of those, do a "same-different" test. Have someone else write down a random sequence of A and B, maybe 20 to 30 steps long, keep it an absolute secret from yourself or whoever is doing the listening, run the same signal into both amps, and have another person plugging and unplugging the speaker cables according to that sequence. So for example, if the first step is B, you listen to it and get a feel for how it sounds, but you don't know that you're listening to A or B. Ready? The person switching disconnects B and then connects the cables to either A or B, whichever is next. You listen; is it the same or is it different? Write down your decision. Then when you're ready, have the person unplug the speaker cables from that amp and plug into the next amp in the sequence, and so on. At the end, compare the same-different decisions with the actual sequence.
Whether an ABX or same-different test, a score of 50% means that you did as well statistically as if you'd tossed a coin to make your decisions.

That would mean that it was not proven that there was an audible difference. That doesn't prove that there isn't an audible difference, but similar results from repeated tests can suggest that.
Yes, it's a lot to go through if you really want to make comparisons. But the bottom line is that you have to compare them at the same volume, and you have to be aware that in sighted tests even the most honest or savvy person can easily perceive sonic differences that may not actually exist. I'm not immune to it either, so I rely on double blind tests when I can.