How a “boring” W⁺W⁻ cross section measurement quietly became one of the sharpest tests of the standard model, and then, almost as an afterthought, set the tightest limits on dimension-6 operators from this final state
A PPC spotlight on the CMS measurement “W⁺W⁻ boson pair production in proton-proton collisions at √s = 13 TeV”, published as Phys. Rev. D 102, 092001 (2020) (arXiv:2009.00119, analysis CMS-SMP-18-004).
There is a certain kind of physics that rarely makes a press release. News item about a cross section that agrees with theory are uncommon. There is no bump, no dark sector, no gathering in the CERN main auditorium. There is just a number, an uncertainty, and a prediction on which that number happens to land.
The production of W⁺W⁻ pairs is exactly that kind of physics. It is also, not coincidentally, the most persistent background in half the searches at the CERN LHC. Looking for H → WW? Congratulations, this is your background. Looking for chargino pairs? Background. Looking for anything with two leptons and missing transverse momentum? You already know. For years the community’s relationship with W⁺W⁻ production has been roughly: please go away, and please be exactly as large as the simulation says you are.
This paper takes the opposite attitude. Rather than subtracting W⁺W⁻ production, it measures it, carefully, twice, using two independent methods. Then, having gone to all that trouble, it uses the leftovers to set the tightest constraints on new electroweak physics anyone had obtained from this final state.
The signature, and why it is hard
The process W⁺W⁻ → ℓ⁺νℓ⁻ν̄ gives two isolated, oppositely charged leptons and two neutrinos. That is all. There is no resonance to fit and no mass peak to point at, the two neutrinos having seen to that. What remains is a shape in the lepton kinematic variables, together with a set of backgrounds that produce very nearly the same thing:
- Drell–Yan production of lepton pairs, which is copious in the same-flavor final state and produces apparent missing transverse momentum only when the detector misbehaves;
- top quark production (tt̄ and tW), which is genuinely two W bosons accompanied by b jets, and therefore differs from the signal only by the jets one manages to tag;
- nonprompt leptons from W+jets events, in which a hadron is misidentified as a lepton;
- and, in smaller amounts, VZ and Wγ* production.
The analysis uses 35.9 fb⁻¹ of proton-proton collision data recorded by CMS in 2016 and attacks this with two complementary strategies, which is the sort of thing one does when one genuinely does not want to be wrong.
The sequential cut analysis is the classical approach. A sequence of requirements is applied to kinematic quantities, with events divided by lepton flavor (different flavor against same flavor) and by jet multiplicity (0 or 1 jet). One clever ingredient, applied throughout, is the projected missing transverse momentum, defined as the component of p⃗Tmiss perpendicular to the nearest lepton, so that a single mismeasured lepton cannot manufacture an apparent imbalance. Top quark events are suppressed by rejecting events that contain a loosely b tagged jet.
The same-flavor final state, where the Drell–Yan contribution is far larger, is then tightened in three cut-based steps: dilepton masses within 15 GeV of the Z boson mass are rejected, the minimum mℓℓ is raised to 40 GeV, and the pTmiss requirement is raised to 55 GeV. Only afterwards does a standard CMS boosted decision tree classifier come into play, to discriminate against the Drell–Yan events that survive all of that. The mass window does the vetoing, and the classifier cleans up behind it. The reward for keeping the 0- and 1-jet categories separate is control over the QCD scale uncertainties that afflict any jet-binned measurement.
The random forest analysis takes the modern route, and the choice of classifier is worth pausing on, because the paper makes a deliberate point of it.
A random forest is not a boosted decision tree. The boosted decision tree, the workhorse of a generation of CMS analyses, grows its trees in sequence, each one trained to correct the mistakes of those before it. A random forest instead grows many binary decision trees independently and in parallel, gives each tree only a random subset of the available features, and then aggregates their verdicts. That randomization is the whole trick. Overfitting committed by any individual tree is averaged away rather than passed down the chain, so the classifier is expected to improve monotonically as trees are added, rather than eventually beginning to degrade. The practical payoff, and the reason this analysis chose it, is that a random forest matches the performance of a boosted decision tree while requiring considerably less tuning of hyperparameters. This analysis is the first in CMS to apply the random forest technique.
Two such classifiers are trained here, one against Drell–Yan production and one against top quark production. Each is fed 14 kinematic and topological features, among them lepton flavor, jet multiplicity, mℓℓ, pTmiss, HT, several azimuthal angle combinations, and the charged-particle recoil. There is no flavor split and no jet split, only purity. Sometimes the best machine learning result is the one that needed the least adjustment.
The two methods have genuinely different weaknesses. The sequential cut analysis retains more top quark contamination, whereas the random forest analysis is more sensitive to the theoretical uncertainties in the W⁺W⁻ transverse momentum spectrum. That is precisely why running both was worth the effort.
The result: a number that lands on the prediction
A simultaneous fit to the four categories of the sequential cut analysis gives
σ(pp → W⁺W⁻) = 117.6 ± 1.4 (stat) ± 5.5 (syst) ± 1.9 (theo) ± 3.2 (lumi) pb = 117.6 ± 6.8 pb.
The NNLO prediction, which includes the gluon fusion contribution and electroweak corrections, is 118.8 ± 3.6 pb. The two central values differ by about 1%. The uncertainty is several times larger than that, so nobody should become emotional about the coincidence, but a 5.8% measurement of a diboson cross section landing on top of a 3% prediction makes for a good day.
The fiducial cross section is, if anything, even closer: 1.529 ± 0.087 pb measured against 1.531 ± 0.043 pb predicted. Split by jet multiplicity, the values are σfid(0 jet) = 1.61 ± 0.10 pb and σfid(1 jet) = 1.35 ± 0.11 pb.
The random forest analysis returns 131.4 ± 8.7 pb, about 12% higher. This is less a contradiction than a measurement of the analyses themselves. By using HT and related observables to suppress the top quark background, the random forest preferentially selects events of low jet multiplicity, and therefore of low W⁺W⁻ transverse momentum, which makes it more sensitive to the theoretical modeling of that spectrum. Taken together, the two numbers say something that the sequential result alone would not.
There is a further measurement worth noting. The 0-jet fiducial cross section is determined as a function of the jet pT threshold, rising from 0.836 to 1.118 pb as that threshold moves from 25 to 60 GeV. This is the experimental version of asking how much the answer depends on where the line was drawn, a question that more measurements could stand to ask out loud.
The most significant plot
The figure that best captures what this analysis could do that others could not is the normalized jet multiplicity distribution. The measurement is possible only because the random forest selection suppresses the top quark background without ever looking at the number of jets.


Figure 7 (Phys. Rev. D 102, 092001 (2020)): Fractions of W⁺W⁻ events with NJ = 0, 1, ≥2 jets. The filled circles show the data after background subtraction and after unfolding for pileup jets and jet energy resolution, and the solid line shows the POWHEG+PYTHIA prediction. The lower panel shows the ratio of the prediction to the measurement.
Why this one? Because it is the plot that only this paper could produce. Every earlier W⁺W⁻ measurement had to impose a jet veto in order to control the top quark background, which makes a measurement of the jet multiplicity somewhere between circular and impossible. Here the top quark rejection is orthogonal to the jet count, so the distribution can simply be measured. It is then unfolded through two response matrices, one for pileup jets and one for jet energy resolution, with no regularization applied.
The measured fractions are 0.773 ± 0.008 ± 0.075, 0.193 ± 0.007 ± 0.043, and 0.034 ± 0.006 ± 0.033 for 0, 1, and at least 2 jets, compared with 0.677, 0.248, and 0.075 from POWHEG. The data prefer fewer jets than the generator does. The uncertainties are large enough that this is a nudge rather than a discovery, but it is a nudge in a direction that matters, because a great many Higgs boson and beyond-the-standard-model analyses depend on how well the simulation models exactly this distribution.
And then, the search for new physics
Having measured everything, the analysis turns the electron-muon mass spectrum into a probe of what is not there. In the language of effective field theory, new heavy physics appears as higher-dimensional operators, here the three CP-conserving dimension-6 operators (𝒪WWW, 𝒪W, 𝒪B) that modify the triple gauge boson vertices. Their contributions grow with energy, so the whole game is played in the last bin, meμ > 1 TeV.
The resulting intervals on the Wilson coefficients are:
| Coefficient | Expected 95% CL (TeV⁻²) | Observed 95% CL (TeV⁻²) |
|---|---|---|
| cWWW/Λ² | [−2.7, 2.7] | [−1.8, 1.8] |
| cW/Λ² | [−5.3, 4.2] | [−3.6, 2.8] |
| cB/Λ² | [−14, 13] | [−9.4, 8.5] |


Figure 10 (Phys. Rev. D 102, 092001 (2020)): Scans of −2Δln L for cWWW/Λ², cW/Λ², and cB/Λ² (left), and the corresponding 68 and 95% CL contours for pairs of coefficients (right), with the 0- and 1-jet categories combined.
The observed intervals are tighter than the expected ones, which sounds suspicious until one looks at the spectrum and finds a modest deficit of events at high meμ, a downward fluctuation that stays within 2 standard deviations of the expectation. With that caveat noted, the numbers stand. These limits were about a factor of 2 more stringent than the contemporaneous ATLAS results and than the previous CMS W⁺W⁻ analysis. For cWWW and cW the sensitivity is comparable to that of the CMS WZ measurement and considerably better for cB, while remaining slightly weaker than the CMS analysis of W⁺W⁻ and WZ production in the lepton plus jets final state.
Not a bad outcome for a measurement of everybody’s least favorite background.
The PPC angle
Guillelmo Gomez-Ceballos was the key analyzer and the driving force behind this measurement. He is the person who kept it moving through the long unglamorous middle, where two independent analyses have to be built, cross-checked against each other, defended in review, and reconciled when they disagree by 12%. There is no shortcut through that work, and no obvious reward at the end of it beyond a number that agrees with theory.
That is rather the point. Production of W⁺W⁻ pairs is the background beneath the Higgs boson measurements, beneath the chargino searches, and beneath a decade of dilepton plus missing momentum analyses at both LHC experiments. Measuring it to 5.8%, and measuring for the first time how its jets are actually distributed rather than assuming it, is infrastructure. It is the sort of result whose value shows up in other people’s papers.
It also fits a pattern this group returns to again and again. A background is never only a background. Measure the thing you were going to subtract, and it tells you something. Here it told us that the standard model is holding up, that our generators put slightly too many jets into W⁺W⁻ events, and that if there is new physics in the triple gauge vertex, it lives above a couple of TeV.
No big press release, no bump, no discovery, but a number and confidence that everyone else needed, measured properly, by someone who did not think twice.
// Christoph Paus / the PPC, MIT
References
- CMS Collaboration, “W⁺W⁻ boson pair production in proton-proton collisions at √s = 13 TeV”, Phys. Rev. D 102 (2020) 092001. DOI: 10.1103/PhysRevD.102.092001. arXiv: 2009.00119. Report nos. CERN-EP-2020-144, CMS-SMP-18-004.
- Public figures and additional material: CMS-SMP-18-004 public results.
- NNLO W⁺W⁻ cross section used as the reference prediction: T. Gehrmann et al., “W⁺W⁻ production at hadron colliders in next-to-next-to-leading order QCD”, Phys. Rev. Lett. 113 (2014) 212001, arXiv:1408.5243.
- NLO QCD corrections to the gluon fusion contribution: F. Caola, K. Melnikov, R. Röntsch, and L. Tancredi, “QCD corrections to W⁺W⁻ production through gluon fusion”, Phys. Lett. B 754 (2016) 275, arXiv:1511.08617.
- Previous CMS W⁺W⁻ measurement at 8 TeV, with limits on anomalous gauge couplings: CMS Collaboration, Eur. Phys. J. C 76 (2016) 401, arXiv:1507.03268.
- ATLAS 13 TeV W⁺W⁻ measurement used for comparison: ATLAS Collaboration, Eur. Phys. J. C 79 (2019) 884, arXiv:1905.04242.
- Effective field theory framework for anomalous triple gauge couplings: C. Degrande et al., “Effective field theory: a modern approach to anomalous couplings”, Annals Phys. 335 (2013) 21, arXiv:1205.4231.
- Dimension-6 operator basis: B. Grzadkowski, M. Iskrzyński, M. Misiak, and J. Rosiek, “Dimension-six terms in the standard model Lagrangian”, JHEP 10 (2010) 085, arXiv:1008.4884; W. Buchmüller and D. Wyler, Nucl. Phys. B 268 (1986) 621.
- Random forest classifiers: L. Breiman, “Random forests”, Machine Learning 45 (2001) 5.
- Why randomizing the features suppresses overfitting: T. K. Ho, “The random subspace method for constructing decision forests”, IEEE Trans. Pattern Anal. Mach. Intell. 20 (1998) 832.
- Random forests compared with other supervised learning methods: R. Caruana and A. Niculescu-Mizil, “An empirical comparison of supervised learning algorithms using different performance metrics”, Proc. 23rd Int. Conf. on Machine Learning (ICML’06) (2006) 161.
Figures reproduced from CMS-SMP-18-004 under the CC-BY-4.0 license. Spotlight prepared for the MIT Particle Physics Collaboration (PPC).
Figures reproduced from CMS-SMP-18-004 under the CC-BY-4.0 license. Spotlight prepared for the MIT Particle Physics Collaboration (PPC).
