Hey, what about noise? Just brainstorming a little... if NONE of the music (music = the non-speech sounds) was being stored, which do you think would be the more likely reason: the music is too complex or the music is too simple? I vote too simple. More realistically then, say that the musical patterns are simple enough that very little about them needs to be stored because they can be easily reconstructed/generated, perhaps with a prompt (e.g., hearing a few seconds or a few fractions of a second of the music, perhaps kinda like having the rule to generate a sequence if given the first term(s) or somesuch bit of variable info).
Now, when we heard the speech, it was mixed in with the music. As far as the 'speech processors' are concerned, wouldn't the music be kinda like noise? (Perhaps I'm stretching the notion of noise, but eh.) The idea is that there's a pattern to the noise, which, once we have it, would allow us to extract the speech message from the signal... which would be much more difficult without knowing that pattern. Does that make any sense?
Okay, assuming that, what does it say about how the signal might be stored? First problem I see: Why can't we recognize the musical pattern in the signal now, as we have it stored? Presumably, we could do so if we heard the signal. So perhaps the signal is stored in such a way that we need the pattern to put the pieces in the correct order. That is, we've stored the signal in bits and pieces here and there, so it's difficult to recall the signal without some key from the music.
So the picture: We don't store the music. We store the speech. But the speech has music as noise mixed in, and it's stored in pieces that need help from some info in the music in order to be reassembled into larger pieces of the original signal.