Implementation and performance evaluation of collective communication with direct GPU-to-GPU communication using the Tightly Coupled Accelerators (TCA) architecture
Implementation and performance evaluation of collective communication with direct GPU-to-GPU communication using the Tightly Coupled Accelerators (TCA) architecture
松本 和也, 塙 敏博, 児玉 祐悦, 藤井 久史, 朴 泰祐: “Implementation and performance evaluation of collective communication with direct GPU-to-GPU communication using the Tightly Coupled Accelerators (TCA) architecture”, IPSJ Transactions on Advanced Computing Systems (ACS), Vol. 8, No. 4, pp. 36-49, 2015. (in Japanese)
IPSJ Transactions on Advanced Computing Systems (ACS)
BiBTeX entry
@article { ACS-8-4:matsumoto,
author = "{松本 和也} and {塙 敏博} and {児玉 祐悦} and {藤井 久史} and {朴 泰祐}",
title = "{密結合並列演算加速機構TCAによるGPU間直接通信におけるCollective通信の実装と性能評価}",
journal = "情報処理学会論文誌コンピューティングシステム (ACS)",
volume = "8",
number = "4",
pages = "36--49",
year = "2015"
}