ENH: load particle definitions from pdg package - #371
Conversation
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
…add caching for PDG particle loading and remove unused dependencies
…g for import from official pdg package
|
I think for particle names they can produce latex. https://pdg.lbl.gov/2026/api/index.html if it is not in They say
there is issue tracker here Dean, and Jurgen could add it at the next release. |
redeboer
left a comment
There was a problem hiding this comment.
Looks good! We just need some more cross-checks and documentation.
| default_particles = load_pdg() | ||
| scikit_hep_particles = load_pdg(source="particle") |
There was a problem hiding this comment.
Could you test whether there are any mismatches between two collections from particle and from pdg? We can now directly use this new feature to find bugs in the particle databases (which are hardcoded).
Note
This is something for a follow-up issue. So just try it first and if you find some problems, post an issue. We don't need that mismatch search to be implemented in the unit tests.
There was a problem hiding this comment.
I have wrote a small script to compare both sources. This is the output:
Details
from collections.abc import Iterable
from qrules.particle import Particle, load_pdg
def index_by_pid(particles: Iterable[Particle]) -> dict[int, Particle]:
"""Index particles by their Monte Carlo particle ID."""
return {particle.pid: particle for particle in particles}
def main() -> None:
"""Report source coverage and particles available only from the PDG API."""
scikit_hep_particles = index_by_pid(load_pdg(source="particle"))
pdg_particles = index_by_pid(load_pdg(source="pdg"))
scikit_hep_ids = set(scikit_hep_particles)
pdg_ids = set(pdg_particles)
only_in_pdg = pdg_ids - scikit_hep_ids
print(f"Scikit-HEP source: {len(scikit_hep_ids)} particles")
print(f"PDG source: {len(pdg_ids)} particles")
print(f"All PDG particles are in Scikit-HEP: {pdg_ids <= scikit_hep_ids}")
print(f"All Scikit-HEP particles are in PDG: {scikit_hep_ids <= pdg_ids}")
print(f"\nParticles only in the PDG source ({len(only_in_pdg)}):")
for pid in sorted(only_in_pdg):
print(f"{pid:>9} {pdg_particles[pid].name}")
if __name__ == "__main__":
main()There was a problem hiding this comment.
Perhaps a variation of that script could be used to update the corresponding database in the particle package? I mean, one that directly gets the data from pdg and then fetches the missing entries for that CSV file.
pdg package
Closes #332
⚙️ Enhancements
.load_pdgnow accepts asourcekeyword-only argument. Withsource="pdg"(default remains"particle"), particle definitions are loaded from the official PDG Python API instead of theparticlepackage provided by scikit-hep, giving a larger and independently sourced set of particle definitions. As of writing, particles loaded this way do not have LaTeX names, because thepdgpackage does not yet support them (particledatagroup/api#42).