My understanding is that the double birefringent layer splits each point of the image into four different points spaced slightly apart. The light then hits the color filters and is filtered appropriately. Since real images are composed of a continuous set of an infinite amount of points, this has the effect of making some of the light from each point fall on multiple sensors instead of just one. In other words, the image is blurred slightly, with the amount of blurring dependent on the amount of separation between each of the four optical points.
For example, if the resolution of an imaging system gives us an airy disk for light exactly equal to the size of each pixel on the sensor, then practically all of the light from some of the points of the image will fall only on one of these pixels. The number of these "special" points will be equal to the number of pixels on the sensor. Points of the image in between these special points will have their light split between two or more pixels. Because our resolution is limited, spatial frequencies higher than our resolution are lost since the sensor can't tell them apart no matter how small the sensor's pixels are.
When the resolution of the imaging system gives us an airy disk significantly smaller than each pixel, then we can come across aliasing since our optical system can separate the points of these higher spatial frequencies but our sensor can't. The sensor records only "discrete" portions of the image, not the continuous range, so these high frequency patterns (higher than the spatial frequency the sensor samples at) manifest as abrupt changes in the resulting sampled image. Hence why shifting to a sensor with a higher number of smaller pixels decreases aliasing.
When you have a color filter matrix, the problem of aliasing becomes much more severe since each pixel will receive light predominantly within a small range of wavelengths and each pixel of the same color is typically separated from other pixels of the same color by at least one pixel. This is the same as having a large spacing between pixels of a monochrome sensor. Light that falls between the pixels is simply lost, so you have a case where the sampling done by the sensor is much lower than the spatial resolution of the optical device.
Having an optical low pass filter helps because it blurs the entire image, effectively increasing the size of the airy disk so that light from some of the areas of the image that is normally lost is instead captured by a pixel. In addition, the higher spatial frequencies are filtered out by this blurring since it increases the amount of separation required between two points of the image to see them as separate points. Splitting the light from each point into four beams means you are literally projecting four different images onto a single sensor, with each image slightly offset from the others.
To my understanding one of the things an OLPF doesn't do is that it doesn't split the light from all points on the image into four beams with each beam landing exactly on one color pixel. A single point at a particular location on the image may have this happen if the system is set up to do so, but if you look at another nearby point you will find that the light is split into four beams with each beam falling on more than one pixel.
That's the way I understand the working of an optical low pass filter. As always, someone correct me if I'm wrong.