1. Limited time only! Sign up for a free 30min personal tutor trial with Chegg Tutors
    Dismiss Notice
Dismiss Notice
Join Physics Forums Today!
The friendliest, high quality science and math community on the planet! Everyone who loves science is here!

PCA and variance on particular axis

  1. Oct 19, 2007 #1
    Hi All:

    If given a set of 3D points data, it's very easy to calculate the covariance matrix and get the principle axises. And the eigenvalue will be the variance on the principle axis. I have a problem that if given a random direction, how do I calculate the variance of the data on the given direction?

    Can anybody help me with this please?

  2. jcsd
  3. Oct 25, 2007 #2

    Chris Hillman

    User Avatar
    Science Advisor

    Speaking of Applicable Geometry

    How much do you know about the relationship between euclidean geometry and mean/variance?

    Your question really concerns the values taken by a quadratic form. Many statistical manipulations (and many useful properties of normal distributions) arise in a geometrically natural manner from manipulating quadratic forms by orthogonal transformations and orthoprojections, together with some notions from affine geometry such as convexity. For example, taking the mean of n variables [itex]x_1, \, x_2, \dots x_n[/itex], where we think of this data as the vector [itex]\vec{x} = x_1 \, \vec{e}_1 + \dots x_n \, \vec{e}_n [/itex], corresponds to taking the orthoprojection (defined using standard euclidean inner product) onto the one dimensional subspace spanned by [itex]\vec{e}_1 + \vec{e}_2 + \dots \vec{e}_n[/itex]. If we adopt a new orthonormal basis including the unit vector [itex]\vec{f}_n = \frac{1}{\sqrt{n}} \, \left( \vec{e}_1 + \vec{e}_2 + \dots \vec{e}_n \right)[/itex], this orthoprojection can be thought of very simply, as simply forgetting all but the last component [itex]\sqrt{n} \, \overline{x}[/itex], which agrees (up to a constant multiple) with the arithmetic mean.

    See M. G. Kendall, A Course in the Geometry of n Dimensions, Dover reprint, and then try the same author's book Multivariate Analysis.

    I must add a caution: do you see why principle component analysis (PCA) is essentially a method for "lying with statistics"? That is, the geometric (or if you prefer, linear algebraic) manipulations of your data set are mathematically valid, but the statistical interpretation is almost always extremely dubious. Fortunately, my remark about the role of euclidean geometry in mathematical statistics holds true for many more legitimate statistical methods, some discussed in the first book by Kendall cited above.
    Last edited: Oct 25, 2007
Share this great discussion with others via Reddit, Google+, Twitter, or Facebook