paperbot · PL 论文追踪

RSS

HiPy: Extracting High-Level Semantics from Python Code for Data Processing

OOPSLA 8(OOPSLA2)2024引用 12
Michael Jungmair, Alexis Engelke, Jana Giceva

尚未生成 AI 速览(可能缺少 API key 或等待下次运行补跑)。

原文摘要(Abstract)

Data science workloads frequently include Python code, but Python’s dynamic nature makes efficient execution hard. Traditional approaches either treat Python as a black box, missing out on optimization potential, or are limited to a narrow domain. However, a deep and efficient integration of user-defined Python code into data processing systems requires extracting the semantics of the entire Python code. In this paper, we propose a novel approach for extracting the high-level semantics by transforming general Python functions into program generators that generate a statically-typed IR when executed. The extracted IR then allows for high-level, domain-specific optimizations and the generation of efficient C++ code. With our prototype implementation, HiPy, we achieve single-threaded speedups of 2–20x for many workloads. Furthermore, HiPy is also capable of accelerating Python code in other domains like numerical data, where it can sometimes even outperform specialized compilers.

链接与引用

DOI 原文 · PDF(开放获取) · DBLP

BibTeX
@article{JungmairEG24,
  title = {HiPy: Extracting High-Level Semantics from Python Code for Data Processing},
  author = {Michael Jungmair and Alexis Engelke and Jana Giceva},
  journal = {Proceedings of the ACM on Programming Languages},
  volume = {8},
  number = {OOPSLA2},
  year = {2024},
  doi = {10.1145/3689737}
}