Group photo

System Software team

Professor Osamu Tatebe

  • Research on parallel and distributed system software for HPC, big data and AI

We conduct research on system software for very large-scale big data analysis, data-intensive computing and high performance computing.

Handling ever larger data requires an I/O mechanism that scales with both the size of the data and the computing speed of the machine. To that end we study parallel and distributed system software that scales out, including distributed file systems, parallel I/O and application frameworks.

Our research uses large-scale PC clusters for HPC, Pegasus (the University of Tsukuba’s supercomputer), Miyabi (operated jointly with the University of Tokyo) and other systems.

On supercomputers, ad hoc file systems — which build a temporary parallel file system out of the local storage of the compute nodes — are seen as a promising way to close the gap between the performance of the parallel file system and that of the computation itself.

We are developing CHFS and FINCHFS as ad hoc file systems. CHFS ranked 21st in the 10-node research category of the June 2023 edition of IO500, the world ranking of I/O performance.

  • Research on caching file systems (CHFS/Cache)

An ad hoc file system is a temporary parallel file system built on node-local storage, and data must therefore be moved between it and the underlying parallel file system. A caching file system is a mechanism that performs this data movement automatically.

We are developing CHFS/Cache, a version of CHFS that works as a caching file system. Conventional caching file systems suffer from poor performance when accessing small files; in CHFS/Cache we address this problem by relaxing the consistency requirements between the cache and the parallel file system.

There was also a problem of degraded performance when flushing modified data back to the parallel file system, which we address by proposing an I/O-aware flushing mechanism.

  • Research on parallel I/O libraries

Ad hoc file systems and similar software are implemented at user level, so some effort is needed before applications can use them. So that applications can be used without modification, we are extending libraries such as MPI-IO, HDF5, Apache Arrow and TensorStore, so that any application built on those libraries works unchanged.

We are also devising a variety of techniques to further improve the I/O performance of applications.

  • Study of the storage system for the next flagship machine

We are studying the storage system for the flagship machine that will succeed Fugaku. The following is an interim report.

  • The Gfarm file system

We research and develop the Gfarm file system as a wide-area distributed file system. Gfarm is a storage system that can be accessed securely over the Internet, can distribute storage across wide areas, scales out in both performance and capacity, has no single point of failure, guarantees data integrity, and can cope with silent data corruption.

Gfarm is in production use in systems such as the HPCI shared storage. The HPCI shared storage is a wide-area storage system used across the domestic HPC infrastructure centred on Fugaku, which is promoted by MEXT; it can be mounted and used from anywhere, including supercomputer centres throughout Japan. Files are automatically replicated to the eastern site (Kashiwa) and the western site (Kobe), so they remain accessible even when a failure occurs.

Everyday life in the team

  • Meetings are held in person. Team meetings are weekly, and the whole-laboratory meeting is roughly monthly.
  • We also hold a reading group within the team. This year we are reading Architecture and Design of the Linux Storage Stack .
  • There are no core hours.
  • On weekday afternoons there are typically two or three people in the laboratory.
  • The laboratory is equipped with a microwave oven, a refrigerator, an electric kettle, a coffee maker, sofas and so on, all of which students are free to use. Snacks and drinks are also on sale on site.
  • The goal for fourth-year undergraduates is to give an oral presentation at the HPC research meeting in March. Of course, if progress is fast, they can present at an international conference.
    • Last year one fourth-year student presented in the poster session of SC24, a top-tier conference.
  • There are a number of enjoyable events, such as the implementation camp and the alumni gathering.
    • The photograph below is from last year’s alumni gathering.

SS team alumni gathering in 2024

Computing resources and operational services of the laboratory

  • We have about seven nodes equipped with Intel Optane persistent memory.
  • We have close to forty nodes with InfiniBand, two of which reach 400 Gbps.
  • Including the nodes used to run services, the team has about seventy of its own nodes, and root is available on all of them.
  • The cluster at the Center for Computational Sciences and the laboratory are connected by a 10G link.

Because we have many experimental nodes on which root privileges are available and even the kernel can be replaced, a wide range of research is possible. When root privileges are not required, the University of Tsukuba’s Cygnus and Pegasus supercomputers are also available, and after discussion with the faculty we sometimes use supercomputers at other universities as well.

The shared file system, DNS and other services run by the team are operated by volunteers within the team, mainly master’s students. Web services and DNS operation for the whole laboratory are handled by the admin group. Neither is compulsory, so you are free either to concentrate on research and produce results, or to gain operational experience alongside your research.

Members

Osamu Tatebe

Osamu Tatebe Professor

  • Distributed file systems
  • Parallel system software

If you are interested in system software, or you want to take on something big, you are very welcome. Work on what you love.

Munenori Maeda

Munenori Maeda Senior Researcher

I am a researcher who came from Fujitsu Limited through an industry-academia collaboration project. Fast distributed data stores are becoming ever more important. Let's work on them together in this lab.

Sohei Koyama

Sohei Koyama D3

I am developing an ad hoc file system. Let's aim for the top of the IO500 together

Kohei Sugihara

Kohei Sugihara D3

  • Node-local burst buffers

This is a lab with more computers than people. If you want to try doing research with clusters and supercomputers, come join us!

Mingzhe Yu

Mingzhe Yu D3

I am thinking about doing research on fault tolerance in distributed learning

Alexander Klassen

Alexander Klassen M2

I am going to do research on optimizing distributed file systems

Motoi Kourakata

Motoi Kourakata M2

Computers are just so much fun

Shingo Hattori

Shingo Hattori M2

I love low-latency, high-bandwidth networks. I want to move data as fast as it can possibly go.

Ryosuke Maeda

Ryosuke Maeda M2

The cries of asynchrony, brutal~

Kogi Yuki

Kogi Yuki M1

In pursuit of the limits of performance.

Hiroki Nishiyama

Hiroki Nishiyama M1

  • Distributed consensus

I'm thinking of doing research related to distributed consensus and redundancy. But if that turns out to be beyond me, I'll work on something else.

Yamada Toshiya

Yamada Toshiya B4

Supercomputers look like fun

Recent Works

  • Fast checkpointing of Large Language Models with TensorStore CHFS
    • Sohei Koyama
    • Kohei Hiraga
    • Osamu Tatebe
    S. Koyama, K. Hiraga and O.Tatebe, “Fast checkpointing of Large Language Models with TensorStore CHFS,” Supercomputing Conference (SC) 23, Poster, Nov. 2023.
  • I/O-Aware Flushing for HPC Caching Filesystem
    • Osamu Tatebe
    • Kohei Hiraga
    • Hiroki Ohtsuji,
    O. Tatebe, K. Hiraga and H. Ohtsuji, "I/O-Aware Flushing for HPC Caching Filesystem," in 2023 IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops), Santa Fe, NM, USA, 2023 pp. 11-17.
  • Accelerating I/O in Distributed Data Processing Systems with Apache Arrow CHFS
    • Sohei Koyama
    • Kohei Hiraga
    • Osamu Tatebe,
    . Koyama, K. Hiraga and O. Tatebe, "Accelerating I/O in Distributed Data Processing Systems with Apache Arrow CHFS," in 2023 IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops), Santa Fe, NM, USA, 2023 pp. 1-4.
  • Cygnus - World First Multihybrid Accelerated Cluster with GPU and FPGA Coupling
    • Boku Taisuke
    • Fujita Norihisa
    • Kobayashi Ryohei
    • Tatebe Osamu
    Taisuke Boku, Norihisa Fujita, Ryohei Kobayashi, and Osamu Tatebe. 2023. Cygnus - World First Multihybrid Accelerated Cluster with GPU and FPGA Coupling. In Workshop Proceedings of the 51st International Conference on Parallel Processing (ICPP Workshops '22). Association for Computing Machinery, New York, NY, USA, Article 8, 1–8. https://doi.org/10.1145/3547276.3548629
    • Sohei Koyama
    • Osamu Tatebe
    S. Koyama and O. Tatebe, "Scalable Data Parallel Distributed Training for Graph Neural Networks," 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), Lyon, France, 2022, pp. 699-707, doi: 10.1109/IPDPSW55747.2022.00121.
    • Osamu Tatebe
    • Hiroki Ohtsuji
    O. Tatebe and H. Ohtsuji, "Caching Support for CHFS Node-local Persistent Memory File System," 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), Lyon, France, 2022, pp. 1103-1110, doi: 10.1109/IPDPSW55747.2022.00182.
    • Osamu Tatebe
    • Kazuki Obata
    • Kohei Hiraga
    • Hiroki Ohtsuji
    Osamu Tatebe, Kazuki Obata, Kohei Hiraga, and Hiroki Ohtsuji. 2022. CHFS: Parallel Consistent Hashing File System for Node-local Persistent Memory. In International Conference on High Performance Computing in Asia-Pacific Region (HPCAsia '22). Association for Computing Machinery, New York, NY, USA, 115–124. https://doi.org/10.1145/3492805.3492807
    • 建部 修見
    • 平賀 弘平
    • 前田 宗則
    • 藤田 典久
    • 小林 諒平
    • 額田 彰
    建部 修見, 平賀 弘平, 前田 宗則, 藤田 典久, 小林 諒平, 額田 彰: “Performance evaluation of the Pegasus big memory supercomputer,” IPSJ SIG Technical Report, 190th HPC Meeting (SWoPP2023), Aug 2023. (in Japanese)
    • 小山 創平
    • 平賀 弘平
    • 建部 修見
    小山 創平, 平賀 弘平, 建部 修見: “I/O acceleration of big data processing with Apache Arrow CHFS,” IPSJ SIG Technical Report, 190th HPC Meeting (HPC190), (in Japanese)
    • 笠井 大暉
    • 建部 修見
    笠井 大暉, 建部 修見: ”Design and implementation of a distributed cache file system,” IPSJ SIG Technical Report, 186th HPC Meeting, 2022-HPC-186, Jul 2022. (in Japanese)
    • 平賀 弘平
    • 建部 修見
    平賀 弘平, 建部 修見: ”MPI-IO/CHFS: Design of MPI-IO for an ad hoc distributed file system utilizing node-local non-volatile memory,” IPSJ SIG Technical Report, 185th HPC Meeting, Vol. 2022-HPC-185, Jul 2022. (in Japanese)
    • 建部 修見
    建部 修見: ”Evaluation of the access performance of the CHFS ad hoc parallel distributed file system,” IPSJ SIG Technical Report on High Performance Computing (HPC), Vol. 2022-HPC-185, No. 31, Jul 2022. (in Japanese)
    • 巨畠 和樹
    • 小山 創平
    • 平賀 弘平
    • 建部 修見
    巨畠 和樹, 小山 創平, 平賀 弘平, 建部 修見: ”A study on the use of node-local storage in exploratory data analysis assuming an HPC environment,” IPSJ SIG Technical Report, 185th HPC Meeting, Vol. 2022-HPC-185, No. 19, Jul 2022. (in Japanese)
    • 巨畠 和樹
    • 建部 修見
    巨畠 和樹, 建部 修見: ”Design of a distributed object storage using non-volatile memory,” IPSJ SIG Technical Report, 184th HPC Meeting, Vol. 2022-HPC-184, No. 3, May 2022. (in Japanese)
    • 平賀 弘平
    • 建部 修見
    平賀 弘平, 建部 修見: "Design of an MPI-IO burst buffer using non-volatile memory on compute nodes,” IPSJ SIG Technical Report, 183rd HPC Meeting, Vol. 2022-HPC0183, No. 24, May 2022. (in Japanese)
    • 建部 修見
    建部 修見: "Design of a cache file system using non-volatile memory on compute nodes,” IPSJ SIG Technical Report, 183rd HPC Meeting, Vol. 2022-HPC-183, No. 8, May 2022. (in Japanese)