There are different, somewhat common, definitions. They are all based on the time it takes from a step input (like the edge of a square wave) for the output to reach its final value.
They are:
[*] Time-Constant: when the output reaches to 37% of its final value
[*] 5 times the Time Constant: the output reaches to <1% of its final value
[*] 10%: the output reaches to <10% of its final value
For audio stuff, I expect that the 10% rule would be adequate but not ideal.
For a Rock Band you may get away with one Time Constant, 37%.
For Classical, or the 1812 Overture, consider 3 to 5 time constants. ( 3 Time constants gets you within 5% of final value. With the huge dynamic range you would want the recovery time to be quite fast to avoid audible distortion.)
Cheers,
Tom
p.s. These numbers are based on somewhat limited experience, so if someone else has different recommendations, believe them!