Results 38 available artifacts 52 functional artifacts 66 reproduced artifacts Cuttlefish: Library for Achieving Energy Efficiency in Multicore Parallel Programs Best Reproducibility Advancement Award Finalist Artifact Enable Simultaneous DNN Services Based on Deterministic Operator Overlap and Precise Latency Prediction Best Reproducibility Advancement Award Finalist Artifact HPAC: Evaluating Approximate Computing Techniques on HPC OpenMP Applications. Best Reproducibility Advancement Award Finalist Artifact Parallel Construction of Module Networks Best Reproducibility Advancement Award Finalist Artifact Reverse-Mode Automatic Differentiation and Optimization of GPU Kernels via Enzyme full strip note Best Reproducibility Advancement Award Finalist Artifact 3D Acoustic-Elastic Coupling with Gravity: The Dynamics of the 2018 Palu, Sulawesi Earthquake and Tsunami Artifact APNN-TC: Accelerating Arbitrary Precision Neural Networks on Ampere GPU Tensor Cores Artifact Accelerating Applications using Edge Tensor Processing Units Artifact Accelerating XOR-based Erasure Coding using Program Optimization Techniques Artifact Accelerating large scale de novo metagenome assembly using GPUs Artifact CAKE: Matrix Multiplication Using Constant-Bandwidth Blocks Artifact Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters Artifact Discovering and Balancing Fundamental Cycles in Large Signed Graphs Artifact Efficient Tensor Core-based GPU Kernels for Structured Sparsity under Reduced Precision Artifact Flare: Flexible In-Network Allreduce Artifact Hardware-supported Remote Persistence for Distributed Persistent Memory Artifact Hybrid, scalable, trace-driven performance modeling of GPGPUs Artifact In-Depth Analyses of Unified Virtual Memory System for GPU Accelerated Computing Artifact Index Launches: Scalable, Flexible Representation of Parallel Task Groups Artifact Krill: A Compiler and Runtime System for Concurrent Graph Processing Artifact LCCG: A Locality-Centric Hardware Accelerator for High Throughput of Concurrent Graph Processing Artifact Minimizing privilege for building HPC containers Artifact On the Parallel I/O Optimality of Linear Algebra Kernels: Near-Optimal Matrix Factorizations Artifact Online Evolutionary Batch Size Orchestration for Scheduling Deep Learning Workloads in GPU Clusters Artifact PAGANI: A Parallel Adaptive GPU Algorithm for Numerical Integration full strip note Artifact PEPPA-X: finding program test inputs to bound silent data corruption vulnerability in HPC applications Artifact Preparing an Incompressible-Flow Fluid Dynamics Code for Exascale-Class Wind Energy Simulations Artifact Productivity, Portability, Performance: Data-Centric Python Artifact Representation of Women in High-Performance Computing Conferences Artifact Ribbon: Cost-Effective and QoS-Aware Deep Learning Model Inference using a Diverse Pool of Cloud Computing Instances Artifact SEEC: Stochastic Escape Express Channel Artifact STM-Multifrontal QR: Streaming Task Mapping Multifrontal QR Factorization Empowered by GCN Artifact Simurgh: A Fully Decentralized and Secure NVMM User Space File System Artifact Systematically Inferring I/O Performance Variability by Examining Repetitive Job Behavior Artifact Tensor processing primitives: a programming abstraction for efficiency and portability in deep learning workloads Artifact The Hidden cost of the Edge: A Performance Comparison ofEdge and Cloud Latencies Artifact cuTS: Scaling Subgraph Isomorphism on Distributed Multi-GPU Systems Using Trie Based Data Structure Artifact ndzip-gpu: Efficient Lossless Compression of Scientific Floating-Point Data on GPUs Artifact Bootstrapping In-situ Workflow Auto-Tuning via Combining Performance Models of Component Applications full strip note Artifact DeltaFS: A Scalable No-Ground-Truth Filesystem For Massively-Parallel Computing Artifact DistGNN: scalable distributed training for large-scale graph neural networks Artifact Dr. Top-k: delegate-centric Top-k on GPUs Artifact Efficient Large-Scale Language Model Training on GPU Clusters Artifact Exploiting User Activeness for Data Retention in HPC Systems Artifact KAISA: An Adaptive Second-order Optimizer Framework for Deep Neural Networks Artifact LMFF: Efficient and Scalable Layered Materials Force Field on Heterogeneous Many-Core Processors Artifact MAPA: Multi-Accelerator Pattern Allocation Policy for Multi-Tenant GPU Servers Artifact Pinpointing Crash-Consistency Bugs in the HPC I/O Stack: A Cross-Layer Approach Artifact SV-Sim: Scalable PGAS-based State Vector Simulation of Quantum Circuits Artifact Temporal Vectorization for Stencils Artifact ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning Artifact Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines A Next-Generation Discontinuous Galerkin Fluid Dynamics Solver with Application to High-Resolution Lung Airflow Simulations Artifact AgEBO-Tabular: Joint Neural Architecture and Hyperparameter Search with Autotuned Data-Parallel Training for Tabular Data full strip note Artifact Clairvoyant Prefetching for Distributed Machine Learning I/O Artifact HatRPC: Hint-Accelerated Thrift RPC over RDMA Artifact High Performance Uncertainty Quantification with Parallelized Multilevel Markov Chain Monte Carlo Artifact LibShalom: Optimizing Small and Irregular-shaped Matrix Multiplications on ARMv8 Multi-Cores Artifact LogECMem: Coupling Erasure-Coded In-Memory Key-Value Stores with Parity Logging Artifact Lunule: An Agile and Judicious Metadata Load Balancer for CephFS Artifact Non-Recurring Engineering (NRE) Best Practices: A Case Study with the NERSC/NVIDIA OpenMP Contract Artifact Online Optimization of File Transfers in High-Speed Networks Artifact Paths to OpenMP in the Kernel Artifact Reducing Redundancy in Data Organization and Arithmetic Calculation for Stencil Computations Artifact Scalable adaptive PDE solvers in arbitrary domains Artifact Understanding, Predicting and Scheduling Serverless Workloads under Partial Interference Artifact Whale: Efficient One-to-Many Data Partitioning in RDMA-assisted Distributed Stream Processing Systems Artifact