← Back Systems — Paper Notes Distributed systems and systems for ML 2014 OSDI OSDI-2014 Scaling Distributed Machine Learning with the Parameter Server 2014-10-06 · distributed-dataflow 2019 NeurIPS NeurIPS-2019 GPipe:Efficient Training of Giant Neural Networks using Pipeline Parallelism 2019-12-06 · pipeline-parallelism 2020 SC SC-2020 ZeRO:Memory Optimizations Toward Training Trillion Parameter Models 2020-11-09 · distributed-dataflow GTC GTC-2020 Megatron-LM:Training Multi-Billion Parameter Language Models Using Model Parallelism 2020-03-13 · model-parallelism 2022 MLSys MLSys-2022 Pathways:Asynchronous Distributed Dataflow for ML 2022-03-23 · distributed-dataflow ← Browse all notes