• TalkBass has been independent since 1998. Add your voice.
    Create a free account to reply to discussions, view embedded media, and browse with fewer display ads.
    Join freeLog in
    Want zero display ads or expanded classifieds tools? Compare plans.

Tonewood Differences Measured: Interesting Study

This is the gist of the audio sample listening statistics:
https://en.wikipedia.org/wiki/Two-alternative_forced_choice#:~:text=Two-alternative forced choice (2AFC,versions of the sensory input.
They would always be present with A0 and then be presented with An and Bn and forced to make a decision of which matched A0. I don't understand how identifying the instrument correctly is a false positive.

If you count a true positive as distinguishing A0 from B0, then saying A0 and A1 are distinct is a false positive. My research career was in machine learning, and ensuring two As were not seen as A and B was an important part of validation. It was hopefully part of the experimental design here.
 
If you considered the 67 people to be a decision making engine to distinguish instances of A from B, then if you just presented A0 and B0, then it wouldn't be considered sufficient if it was machine learning. So presenting A0...An and B0...Bn and seeing if A0..An end up in the A bucket and B0..n in the B bucket and with what reliability is what you need. In machine learning if it wasn't consistent you might throw more data at the training set, provided you didn't overfit. But if some A0..n end up seen as B and vice versa to a significant degree it suggests that given the dimensions in the data there isn't a clear delineation. If so, then it changes what you need for statistical significance.
 
If you considered the 67 people to be a decision making engine to distinguish instances of A from B, then if you just presented A0 and B0, then it wouldn't be considered sufficient if it was machine learning. So presenting A0...An and B0...Bn and seeing if A0..An end up in the A bucket and B0..n in the B bucket and with what reliability is what you need. In machine learning if it wasn't consistent you might throw more data at the training set, provided you didn't overfit. But if some A0..n end up seen as B and vice versa to a significant degree it suggests that given the dimensions in the data there isn't a clear delineation. If so, then it changes what you need for statistical significance.
I look at it more from statistics in medicine. A test is designed to distinguish identify a disease. So that could be that you do a blood test where a certain value is set as positive and all numbers below that are negatives. A false negative is any negative lab result where the disease actually exists. A false positive is any positive test where the disease doesn't exist.

This test was setup to see if the listener could correctly identify the type of wood listening to one baseline recording and then two random recordings where one was the same wood and the other was not. The way I see it, a false positive and false negative don't exist in this type of 2 alternative forced choice testing. There just isn't a positive or negative result to the test. Either the listener correctly picks the correct sample or they incorrectly pick an incorrect sample. You can't put all the items in the test simultaneously, that isn't how the test is designed. Statistically this type of testing is considered significant when the ability to correctly identify the answer is 75%.
 
  • Like
Reactions: TyBo
Enough with physics and statistics.
Here is a real test from years ago conducted here. Don't cheat and see how you fare!
SCRAP LUMBER BASS vs. ALDER BASS - Can You Tell The Difference??
This sort of testing gets thrown up here and is completely useless. I have no idea what the different basses sound like, so how can I guess which is which? The study cited does the more correct thing which is to give a sample of each and then ask if you can tell which is which with a different recording.

Also, the links don't work for the audio files so it is even less useful.

Edit:
This is like me throwing up 3 white colors and asking who here can tell me which is 1) Chantilly Lace 2) Decorators White and 3) Simply White. And then when people just randomly select white color names saying that they are the same color or can't be distinguished in an actual room.
 
Last edited:
  • Like
Reactions: -Asdfgh-
Plywood is used in double basses as a tone wood.

Plywood is used in a lot of string instruments. I wouldn't consider it a tone wood per se, because what woods are laminated together in any particular sheet of plywood could be pretty much anything. It's used as a cost savings, although over the years some people have developed a love of the sound of particular plywood basses, like Kays and American Standards.

However, pretty much any DB player would tell you that the sound of even low-end carved instruments is audibly superior to the best plywoods. Even Kay owners would probably agree, but they use Kays because of the sound it produces is different than what a carved bass produces.

Of course, there has never been much argument about the importance of the woods used in string instruments with the exception of electric string instruments.
 
  • Like
Reactions: James Collins
Plywood is used in a lot of string instruments. I wouldn't consider it a tone wood per se, because what woods are laminated together in any particular sheet of plywood could be pretty much anything. It's used as a cost savings, although over the years some people have developed a love of the sound of particular plywood basses, like Kays and American Standards.

However, pretty much any DB player would tell you that the sound of even low-end carved instruments is audibly superior to the best plywoods. Even Kay owners would probably agree, but they use Kays because of the sound it produces is different than what a carved bass produces.

Of course, there has never been much argument about the importance of the woods used in string instruments with the exception of electric string instruments.
I never said it was a good tonewood, but some people prefer it.
 
I look at it more from statistics in medicine. A test is designed to distinguish identify a disease. So that could be that you do a blood test where a certain value is set as positive and all numbers below that are negatives. A false negative is any negative lab result where the disease actually exists. A false positive is any positive test where the disease doesn't exist.

It's correct for the experiment as done.

This test was setup to see if the listener could correctly identify the type of wood listening to one baseline recording and then two random recordings where one was the same wood and the other was not. The way I see it, a false positive and false negative don't exist in this type of 2 alternative forced choice testing.

It just strikes me as incomplete experimental design when testing comparison of Ai to Aj of a set A0...An and ditto for B is easy to add and can weed out systematic issues elsewhere. It's a good test, and is used in areas such as psychology research for that reason. E.g., if people thing A7 is different to all of A0..An and you've used A7 in the rest of the testing, it may skew the results.
 
This sort of testing gets thrown up here and is completely useless. I have no idea what the different basses sound like, so how can I guess which is which? The study cited does the more correct thing which is to give a sample of each and then ask if you can tell which is which with a different recording.

Also, the links don't work for the audio files so it is even less useful.

Edit:
This is like me throwing up 3 white colors and asking who here can tell me which is 1) Chantilly Lace 2) Decorators White and 3) Simply White. And then when people just randomly select white color names saying that they are the same color or can't be distinguished in an actual room.

It does probe whether people who say they know for a fact an X bass with Y bits sounds like can do so reliably...

I've considered putting up a test in which I claim I have recorded three basses and for people to say which is why and simply record the same bass three times played in three different ways. But I am no good at being that deceitful, even for 'science'.
 
The problems with the scrap lumber tests were many.

For starters, look at the poll options:


  • Clips X and Y are LUMBER, Clip Z is ALDER
    40 vote(s)
    13.9%

  • Clips X and Y are ALDER, Clip Z is LUMBER
    37 vote(s)
    12.8%

  • Clip X is ALDER, Clips Y and Z are LUMBER
    66 vote(s)
    22.9%

  • I can't hear any difference!
    145 vote(s)
    50.3%

    The recordings themselves were all mp3s which have reduced fidelity.

    How many people listened over a great sound system? Most people probably were listening with earbuds, computer speakers, etc. so the subtleties of the recordings would be lost.

    The whole premise of the poll is whether or not people can identify alder over scrap lumber. 50% of the responses couldn't hear any difference while choice #3 got the next highest percentage.

    So the results seem to support the thesis that people in fact cannot easily identify what type of wood was used. But what makes any of the respondents experts in the sounds of particular tonewoods. Most responses were probably based on "alder should sound better, so I picked the best sounding clip". The people who did get it right most likely just made a lucky choice :whistle:

    With less than 300 votes the margin of error in the poll is pretty high. 14%, 13%, 23% (all rounded) shows that while the remaining 50% heard a difference, it was largely a coin toss what people thought they had heard. It's the folks that couldn't tell any difference that seems most important, because it supports the idea that there is no difference! If half heard no difference at all and less than a quarter could correctly ID which wood was whcih leads to a conclusion that seems to be the real point of the poll: noone can hear any difference between woods. Ergo, wood doesn't matter, it's only the pickups and tone controls that matter.

    But if you remove the actual ID of which wood is which, half of the people did hear a difference although we know the woods were different from the very beginning so there could be confirmation bias. Still, considering the limitations of the fidelity of what people actually were listening to, the fact that half could hear a difference is more signficant than the fact that half couldn't.
 
I never said it was a good tonewood, but some people prefer it.

I was referring to the fact that it's not a specific wood, it's more like a class, the same way rosewood can be Brazilian, Indian and other varieties. Baltic birch plywood is used for speaker cabinets, it's very rigid. I doubt that the plywood chosen for a Kay bass was birch, more likely spruce or maple plywood and certainly not as many plys as for speaker cabinets because rigidity is the last thing you want, right?
 
  • Like
Reactions: -Asdfgh-
The problems with the scrap lumber tests were many.

For starters, look at the poll options:


  • Clips X and Y are LUMBER, Clip Z is ALDER
    40 vote(s)
    13.9%

  • Clips X and Y are ALDER, Clip Z is LUMBER
    37 vote(s)
    12.8%

  • Clip X is ALDER, Clips Y and Z are LUMBER
    66 vote(s)
    22.9%

  • I can't hear any difference!
    145 vote(s)
    50.3%

    The recordings themselves were all mp3s which have reduced fidelity.

    How many people listened over a great sound system? Most people probably were listening with earbuds, computer speakers, etc. so the subtleties of the recordings would be lost.

    The whole premise of the poll is whether or not people can identify alder over scrap lumber. 50% of the responses couldn't hear any difference while choice #3 got the next highest percentage.

    So the results seem to support the thesis that people in fact cannot easily identify what type of wood was used. But what makes any of the respondents experts in the sounds of particular tonewoods. Most responses were probably based on "alder should sound better, so I picked the best sounding clip". The people who did get it right most likely just made a lucky choice :whistle:

    With less than 300 votes the margin of error in the poll is pretty high. 14%, 13%, 23% (all rounded) shows that while the remaining 50% heard a difference, it was largely a coin toss what people thought they had heard. It's the folks that couldn't tell any difference that seems most important, because it supports the idea that there is no difference! If half heard no difference at all and less than a quarter could correctly ID which wood was whcih leads to a conclusion that seems to be the real point of the poll: noone can hear any difference between woods. Ergo, wood doesn't matter, it's only the pickups and tone controls that matter.

    But if you remove the actual ID of which wood is which, half of the people did hear a difference although we know the woods were different from the very beginning so there could be confirmation bias. Still, considering the limitations of the fidelity of what people actually were listening to, the fact that half could hear a difference is more signficant than the fact that half couldn't.

What bitrate were the mp3s at? If 320kpbs, then there's no source fidelity issue based on past testing showing even professed audiophiles can't distinguish between that and uncompressed audio in pristine reproduction settings. You're just left with factors such as reproduction fidelity, hearing fidelity. And I would, as you did, divide the results as 50/50. 50/50 means no statistical ability to tell the difference, but with the noted caveats.
 
The problems with the scrap lumber tests were many.

For starters, look at the poll options:


  • Clips X and Y are LUMBER, Clip Z is ALDER
    40 vote(s)
    13.9%

  • Clips X and Y are ALDER, Clip Z is LUMBER
    37 vote(s)
    12.8%

  • Clip X is ALDER, Clips Y and Z are LUMBER
    66 vote(s)
    22.9%

  • I can't hear any difference!
    145 vote(s)
    50.3%

    The recordings themselves were all mp3s which have reduced fidelity.

    How many people listened over a great sound system? Most people probably were listening with earbuds, computer speakers, etc. so the subtleties of the recordings would be lost.

    The whole premise of the poll is whether or not people can identify alder over scrap lumber. 50% of the responses couldn't hear any difference while choice #3 got the next highest percentage.

    So the results seem to support the thesis that people in fact cannot easily identify what type of wood was used. But what makes any of the respondents experts in the sounds of particular tonewoods. Most responses were probably based on "alder should sound better, so I picked the best sounding clip". The people who did get it right most likely just made a lucky choice :whistle:

    With less than 300 votes the margin of error in the poll is pretty high. 14%, 13%, 23% (all rounded) shows that while the remaining 50% heard a difference, it was largely a coin toss what people thought they had heard. It's the folks that couldn't tell any difference that seems most important, because it supports the idea that there is no difference! If half heard no difference at all and less than a quarter could correctly ID which wood was whcih leads to a conclusion that seems to be the real point of the poll: noone can hear any difference between woods. Ergo, wood doesn't matter, it's only the pickups and tone controls that matter.

    But if you remove the actual ID of which wood is which, half of the people did hear a difference although we know the woods were different from the very beginning so there could be confirmation bias. Still, considering the limitations of the fidelity of what people actually were listening to, the fact that half could hear a difference is more signficant than the fact that half couldn't.

What bitrate were the mp3s at? If 320kpbs, then there's no source fidelity issue based on past testing showing even professed audiophiles can't distinguish between that and uncompressed audio in pristine reproduction settings. You're just left with factors such as reproduction fidelity, hearing fidelity. And I would, as you did, divide the results as 50/50. 50/50 means no statistical ability to tell the difference, but with the noted caveats.
 
The problems with the scrap lumber tests were many.

For starters, look at the poll options:


  • Clips X and Y are LUMBER, Clip Z is ALDER
    40 vote(s)
    13.9%

  • Clips X and Y are ALDER, Clip Z is LUMBER
    37 vote(s)
    12.8%

  • Clip X is ALDER, Clips Y and Z are LUMBER
    66 vote(s)
    22.9%

  • I can't hear any difference!
    145 vote(s)
    50.3%

    The recordings themselves were all mp3s which have reduced fidelity.

    How many people listened over a great sound system? Most people probably were listening with earbuds, computer speakers, etc. so the subtleties of the recordings would be lost.

    The whole premise of the poll is whether or not people can identify alder over scrap lumber. 50% of the responses couldn't hear any difference while choice #3 got the next highest percentage.

    So the results seem to support the thesis that people in fact cannot easily identify what type of wood was used. But what makes any of the respondents experts in the sounds of particular tonewoods. Most responses were probably based on "alder should sound better, so I picked the best sounding clip". The people who did get it right most likely just made a lucky choice :whistle:

    With less than 300 votes the margin of error in the poll is pretty high. 14%, 13%, 23% (all rounded) shows that while the remaining 50% heard a difference, it was largely a coin toss what people thought they had heard. It's the folks that couldn't tell any difference that seems most important, because it supports the idea that there is no difference! If half heard no difference at all and less than a quarter could correctly ID which wood was whcih leads to a conclusion that seems to be the real point of the poll: noone can hear any difference between woods. Ergo, wood doesn't matter, it's only the pickups and tone controls that matter.

    But if you remove the actual ID of which wood is which, half of the people did hear a difference although we know the woods were different from the very beginning so there could be confirmation bias. Still, considering the limitations of the fidelity of what people actually were listening to, the fact that half could hear a difference is more signficant than the fact that half couldn't.

What bitrate were the mp3s at? If 320kpbs, then there's no source fidelity issue based on past testing showing even professed audiophiles can't distinguish between that and uncompressed audio in pristine reproduction settings. You're just left with factors such as reproduction fidelity, hearing fidelity. And I would, as you did, divide the results as 50/50. 50/50 means no statistical ability to tell the difference, but with the noted caveats.
 
If each listener has enough data points, both inter and intra-observer variability can be accurately measured. That would contribute a lot towards the "some people can hear it and others cannot" part of the equation.

Due to the fact that hearing in humans is variable for many reasons, it's axiomatic that some people will be able to hear diffenences and some people won't. Removing the randomness of that variable would be daunting. Also whether the people that can't hear differeneces is due to the lack of a differences or the listeners hearing issues could easily skew the results.
 
It does probe whether people who say they know for a fact an X bass with Y bits sounds like can do so reliably...

I've considered putting up a test in which I claim I have recorded three basses and for people to say which is why and simply record the same bass three times played in three different ways. But I am no good at being that deceitful, even for 'science'.
I get the joke I think, but as far as science is concerned i could see the psychologic problem with telling people they expect to hear a difference and it creating a difference in 3 identical tracks. It would be fun, but would not really prove anything. Maybe this is part of the difference between testing humans and machines.
 
Due to the fact that hearing in humans is variable for many reasons, it's axiomatic that some people will be able to hear diffenences and some people won't. Removing the randomness of that variable would be daunting. Also whether the people that can't hear differeneces is due to the lack of a differences or the listeners hearing issues could easily skew the results.

Having enough listeners is the only way to partly control for that, and the possibility that 1 in a million people can hear a difference will never be eliminated. It's just like a drug trial...people's immune systems differ, and the 2 tails of the bell curve will differ in the extreme.
 
Sorry that the links no longer work. You had to hear it for yourself to really be the judge but a lot of people though they were remarkably similar and for many, so much that one couldn't be sure.
I was a recording engineer in big studios for many years.
No one (with the possible exception of the player), could have ever told the difference in a mix.

Edit:
This is like me throwing up 3 white colors and asking who here can tell me which is 1) Chantilly Lace 2) Decorators White and 3) Simply White. And then when people just randomly select white color names saying that they are the same color or can't be distinguished in an actual room.

I don't view it that way.
I say who cares what color white it is. Can you prefer the same one consistently?

And to sampling rate:
What difference does it make? It's a comparative test and all are tested equally.
Since you can't hear it, you have no idea of the fidelity from 2011.

BTW, Dan Atkins was a really good luthier and helped a lot of people out with set up and repairs.
He also sold high quality basses on this site at really fair prices.
 

Latest posts