Major upgrades successfully completed

Our center is finally fully back online after a period of major upgrades, which put our center into some kind of hibernation mode. During the downtime, we upgraded our infrastructure by replacing an old chiller with a completely new, more efficient one—bringing us closer to our goal of becoming a carbon-neutral site. The Chiller update started on January 21 and was scheduled for three weeks, but due to installation issues had to be prolonged by another week. The startup after the longest shutdown in many years also took longer, because many servers did not come back immediately, which caused unexpected delays. The swift action of the CMS Computing Operations team prevented major production delays by keeping a tight focus on recovering processing that relied on MIT based data. Our center started to process the urgent requests on Februay 25 and is full back online since this weekend (March 8/9).

While in hibernation mode we had limited power left in a few racks and we used the time to deploy a new much more powerful CephFS storage system, which now supports a large fraction of the CMS operations. All servers within CephFS are housed in UPS-backed racks and are connected via 100 Gbs fibers, ensuring high availability and performance. However, due to space constraints, CephFS cannot fully replace our original mass storage (HDFS) at this time, so we will continue relying on HDFS for specific needs.

Leave a Comment