Ganglia Monitoring System



(Back to docs.huihoo.com)

Introduction

Ganglia is a scalable distributed monitoring system for high-performance computing systems such as clusters and Grids. It is based on a hierarchical design targeted at federations of clusters. It leverages widely used technologies such as XML for data representation, XDR for compact, portable data transport, and RRDtool for data storage and visualization. It uses carefully engineered data structures and algorithms to achieve very low per-node overheads and high concurrency. The implementation is robust, has been ported to an extensive set of operating systems and processor architectures, and is currently in use on thousands of clusters around the world. It has been used to link clusters across university campuses and around the world and can scale to handle clusters with 2000 nodes.

Documents

• Monitoring Temperature and Fan Speed Using Ganglia and Winbond Chips (2006)
• The ganglia distributed monitoring system: design, implementation, and experience (2004)

Links

• http://www.ganglia.info/
• http://ganglia.sourceforge.net/
• UC Berkeley Millennium Demo
• Grids and Clusters Group Demo
• http://download.huihoo.com/ganglia/