Electrostatic scaling laws across the proteome — need expert feedback

Join the discussion
Ask a follow-up here, or get your own question answered by working scientists, mathematicians and engineers — people, not an autocomplete.
Real named experts · corrections over time · the nuance an AI answer skips
7 replies · 795 views
Planobilly
Messages
446
Reaction score
113
TL;DR
Computed dipole/quadrupole for all 550K AlphaFold proteins. Average scaling matches random prediction, but functional families show three distinct regimes. Helicase charge suppression survives composition controls. Zenodo DOI included. Seeking critique.
I want to be upfront: I don't have academic credentials in biology or bioinformatics. My background is in physics and engineering — I hold several patents and spent my career solving applied problems. I'm 78 years old, retired, and I run a small private AI lab out of my home in Florida purely for my own curiosity. The centerpiece is an NVIDIA DGX Spark (128 GB unified memory, Grace Blackwell architecture), which I use alongside local LLMs and custom code to explore whatever catches my attention. I also run a small independent news analysis site. None of this is funded or affiliated with any institution.

The Spark gave me the compute to do something I wouldn't have attempted otherwise. Over the past few months I wrote a custom C# electrostatics pipeline and ran it against all 550,122 proteins in the AlphaFold/SwissProt database (version 4) — computing dipole moments, quadrupole tensors, and solvation energies using AMBER ff14SB partial charges, Henderson-Hasselbalch protonation at pH 7.4, and OBC-II Generalized Born solvation. The full run completed with zero computational failures.

Here's what I found, and why I think it's worth someone with domain expertise taking a look.


The Global Result
The proteome-wide average dipole moment scales as N0.831. The random compact-sphere prediction is N5/6 = 0.833. A match to within 0.002.
That means the average protein's charge placement looks statistically random. But the average hides something.


Function-Specific Deviations
Different protein families organize their charges in fundamentally different ways that correspond to what they do. When I broke the dataset into 22 functional categories and ran ANCOVA, the likelihood ratio rejected a common scaling exponent decisively (χ² = 13,724 for quadrupole, p effectively zero). Three distinct regimes emerged:

RegimeFamiliesDipole Exp.Interpretation
SuperlinearKinases (1.10), Motors (1.09), Immunoglobulins (1.17), Receptors (1.00)> 0.9Charges amplify with size — long-range electrostatic steering
ModerateProteases (0.90), Polymerases (0.85), Chaperones (0.79)0.65 – 0.9Near the random baseline — electrostatics not the primary constraint
SuppressedHelicases (0.57), Cytochromes (0.64), Oxidoreductases (0.62)< 0.65Active charge cancellation — the opposite of random

The Strongest Evidence: Suppression
To rule out amino acid composition as the explanation, I ran a spatial null model: for 10,097 proteins, I shuffled partial charges across atomic positions 100 times each, preserving composition but randomizing spatial arrangement.

Helicases have a real exponent of 0.44 versus a shuffled expectation of 0.94. That gap of −0.50 means their charges are specifically arranged to cancel polarity. Composition can't explain it. Shape alone can't explain it (I controlled for radius of gyration).


Experimental Validation
I compared AlphaFold-predicted structures against 19,104 X-ray crystal structures (resolution ≤ 2.5 Å):

  • Solvation energy correlation: r = 0.97
  • Quadrupole: r = 0.79
  • Dipole: r = 0.73
The functional ANOVA on experimental structures alone still holds (quadrupole F = 18.8, p < 10−54).

One Unexpected Result
Ribosomal proteins show a quadrupole exponent of essentially zero (0.09) — flat. Their charge architecture appears frozen, consistent with their status as some of the oldest conserved structures in biology, roughly 3.5 billion years old.


The full paper is published on Zenodo:

DOI: 10.5281/zenodo.20411754
I also wrote a plain-language version on Medium for anyone who wants the shorter explanation.

I'm not claiming this is groundbreaking or that I've discovered something the field has missed. I had the compute, I had a question, and these are the results. What I'd like to know:

  1. Does the methodology hold up? Are there obvious flaws in the electrostatics pipeline or the statistical approach?
  2. Is the composition-preserving shuffle a valid null model, or is there a better control?
  3. Has anyone seen similar functional clustering by electrostatic scaling in the literature? I searched but may have missed prior work.
  4. Are there alternative explanations for the helicase suppression result that I haven't considered?
I appreciate any feedback, especially if it's critical. I'd rather find out I'm wrong here than find out later.

— Billy Simmons
 
Biology news on Phys.org
I'll be checking back for responses and happy to answer questions or provide additional data. Fair warning — Wednesday morning I'm getting on my boat and heading to the Bahamas for four or five days, so if I'm slow to reply after that, I'm not ignoring you.

You're welcome to come along, either in person or vicariously through the internet. I can send back some pictures and analysis of the marine biology I encounter on the trip. If there's interest, I'll post them here. If anyone wants to reach me directly, my inbox here on PF is open.

Cheers,

Billy
 
Planobilly said:
TL;DR: Computed dipole/quadrupole for all 550K AlphaFold proteins. Average scaling matches random prediction, but functional families show three distinct regimes. Helicase charge suppression survives composition controls. Zenodo DOI included. Seeking critique.
Seeking feedback on a personal theory is disallowed by the forum rules you agreed to when you joined. From the rules:

Speculative or Personal Theories:
Physics Forums is not intended as an alternative to the usual professional venues for discussion and review of new ideas, e.g. personal contacts, conferences, and peer review before publication. If you have a new theory or idea, this is not the place to look for feedback on it or to ask for help developing and publishing it.

And note that your Zenodo reference is self-published, whereas acceptable references here must appear in credible, peer-reviewed publications.
 
Thanks for the feedback. I'm not sure this qualifies as a personal theory. This is a computational result — I ran all 550,122 proteins in the AlphaFold/SwissProt database through an electrostatics pipeline using standard AMBER ff14SB partial charges and reported what the numbers showed. That's data, not speculation. I didn't propose a new theory of physics. I measured something and asked if the methodology holds up.

Cheers,

Billy
 
Reply
  • Like
Likes   Reactions: BillTre
Planobilly said:
That's data, not speculation. I didn't propose a new theory of physics. I measured something and asked if the methodology holds up.
Seeing if the "methodology holds up" is exactly what the referees agree to do when a journal evaluates submitted work. Get your paper published and then we can talk about it here.
 
I'm not clear what the "N"s mean here:
Planobilly said:
N0.831. The random compact-sphere prediction is N5/6 = 0.833
Some kind of dipole number I guess.

General comments:
Biochemistry texts cover some of this. Proteins use changes on amino acids in different ways.
Helices will have to cancel charges to a degree to be in a helix. Helices going through membranes will have charges away from the membrane contacting parts in stably be in the membrane. Channel proteins (through membranes) will have their charges by the channel and away from the membrane contacting parts.

Charges can be used in other proteins to attract and bind other particular molecules for various reasons.

Charges are often on the outside of cytoplasmic proteins in contact with the hydrophilic surroundings. Some can be paired up (plus and minus) internally to strengthen the structure.
 
After a Mentor discussion, the thread is reopened provisionally (see below).

Planobilly said:
The Spark gave me the compute to do something I wouldn't have attempted otherwise. Over the past few months I wrote a custom C# electrostatics pipeline and ran it against all 550,122 proteins in the AlphaFold/SwissProt database (version 4) — computing dipole moments, quadrupole tensors, and solvation energies using AMBER ff14SB partial charges, Henderson-Hasselbalch protonation at pH 7.4, and OBC-II Generalized Born solvation. The full run completed with zero computational failures.
@Planobilly -- Can you give us a reference to where you got the equations that you implemented in your code? Have they been published in textbooks, or been published in the peer-reviewed literature? Thanks.
 
Reply
  • Like
Likes   Reactions: Dale