Problem-solving in High Performance Computing: A Situational Awareness Approach with Linux (Paperback)
暫譯: 高效能運算中的問題解決:基於情境意識的 Linux 方法 (平裝本)

Igor Ljubuncic

買這商品的人也買了...

商品描述

Problem-Solving in High Performance Computing: A Situational Awareness Approach with Linux focuses on understanding giant computing grids as cohesive systems. Unlike other titles on general problem-solving or system administration, this book offers a cohesive approach to complex, layered environments, highlighting the difference between standalone system troubleshooting and complex problem-solving in large, mission critical environments, and addressing the pitfalls of information overload, micro, and macro symptoms, also including methods for managing problems in large computing ecosystems.

The authors offer perspective gained from years of developing Intel-based systems that lead the industry in the number of hosts, software tools, and licenses used in chip design. The book offers unique, real-life examples that emphasize the magnitude and operational complexity of high performance computer systems.

  • Provides insider perspectives on challenges in high performance environments with thousands of servers, millions of cores, distributed data centers, and petabytes of shared data
  • Covers analysis, troubleshooting, and system optimization, from initial diagnostics to deep dives into kernel crash dumps
  • Presents macro principles that appeal to a wide range of users and various real-life, complex problems
  • Includes examples from 24/7 mission-critical environments with specific HPC operational constraints

商品描述(中文翻譯)

高效能計算中的問題解決:以 Linux 為基礎的情境意識方法專注於理解大型計算網格作為一個整體系統。與其他關於一般問題解決或系統管理的書籍不同,本書提供了一種針對複雜、分層環境的整體方法,強調獨立系統故障排除與在大型、任務關鍵環境中進行複雜問題解決之間的差異,並解決資訊過載、微觀和宏觀症狀的陷阱,同時包括在大型計算生態系統中管理問題的方法。

作者提供了多年開發基於 Intel 的系統所獲得的觀點,這些系統在芯片設計中使用的主機數量、軟體工具和許可證數量上領先業界。本書提供了獨特的真實案例,強調高效能計算機系統的規模和操作複雜性。

- 提供對高效能環境中面臨的挑戰的內部觀點,這些環境擁有數千台伺服器、數百萬個核心、分散式數據中心和數PB的共享數據
- 涵蓋從初步診斷到深入分析內核崩潰轉儲的分析、故障排除和系統優化
- 提出適用於廣泛用戶和各種真實複雜問題的宏觀原則
- 包含來自 24/7 任務關鍵環境的例子,並考慮特定的 HPC 操作限制