My understanding was that
[tex]\partial^\mu \phi = g^{\mu\nu}\partial_\nu \phi[/tex]
is used to define the contravariant derivative [tex]\nabla^{\mu}[/tex] ([tex]\nabla[/tex] rather than [tex]\partial[/tex] since we're in GR, though as you've said it reduces to [tex]\partial[/tex] in the case of a scalar function).
Given a differentiable manifold, we can consider one-form fields induced upon the manifold by scalar fields: given a scalar field [tex]\phi[/tex], we get a one-form field with components [tex]\partial_\mu \phi[/tex]. We can also consider vector fields induced by curves through the manifold: given a curve [tex]\gamma: \lambda \in \Re \rightarrow p\in M[/tex], we get a vector field with components [tex]\frac{\partial x^\mu}{\partial \lambda}[/tex].
From this starting point, tensor products can be used to build tensor fields of higher valences (so in addition to the (0,1) fields - the one-forms - and the (1,0) fields - the vectors - we can get tensor fields of valence (m,n)). We then select some particular (0,2) tensor field, and decide that this shall be our metric g. g will take two vectors as arguments, and deliver a scalar. That is,
[tex]g(X,Y) = \chi[/tex]
or if you prefer,
[tex]g_{\mu\nu} X^\mu Y^\nu = \chi[/tex]
But that means that [tex]g_{\mu\nu}X^\mu[/tex] has, in effect, an empty argument place which could be filled by a vector; i.e. it is something which will map vectors to scalars - in other words, a one-form. So the notation [tex]X_\nu[/tex] is introduced as shorthand for [tex]g_{\mu\nu}X^\mu[/tex].
The same trick, using the inverse of the metric (i.e. the (2,0) tensor field such that [tex]g_{\mu\nu}g^{\nu\rho} = \delta^{\rho}_{\mu}[/tex]) will allow you to link any one-form (components [tex]p_{\mu}[/tex]) with a particular vector (components [tex]p^\mu[/tex]). In particular, the one-form field with components [tex]\partial_\mu \phi[/tex] has an associated vector field [tex]\partial^\mu \phi[/tex], defined by
[tex]\partial^\mu \phi = g^{\mu\nu}\partial_\nu \phi[/tex]