类似 BLAS 的库实例化软件框架
Recipient of the 2023 James H. Wilkinson Prize for Numerical Software
Recipient of the 2020 SIAM Activity Group on Supercomputing Best Paper Prize
BLIS is an award-winning portable software framework for instantiating high-performance BLAS-like dense linear algebra libraries. The framework was designed to isolate essential kernels of computation that, when optimized, immediately enable optimized implementations of most of its commonly used and computationally intensive operations. BLIS is written in ISO C99 and available under a new/modified/3-clause BSD license. While BLIS exports a new BLAS-like API, it also includes a BLAS compatibility layer which gives application developers access to BLIS implementations via traditional BLAS routine calls. An object-based API unique to BLIS is also available.
For a thorough presentation of our framework, please read our ACM Transactions on Mathematical Software (TOMS) journal article, "BLIS: A Framework for Rapidly Instantiating BLAS Functionality". For those who just want an executive summary, please see the Key Features section below.
In a follow-up article (also in ACM TOMS), "The BLIS Framework: Experiments in Portability", we investigate using BLIS to instantiate level-3 BLAS implementations on a variety of general-purpose, low-power, and multicore architectures.
An IPDPS'14 conference paper titled "Anatomy of High-Performance Many-Threaded Matrix Multiplication" systematically explores the opportunities for parallelism within the five loops that BLIS exposes in its matrix multiplication algorithm.
For other papers related to BLIS, please see the Citations section below.
It is our belief that BLIS offers substantial benefits in productivity when compared to conventional approaches to developing BLAS libraries, as well as a much-needed refinement of the BLAS interface, and thus constitutes a major advance in dense linear algebra computation. While BLIS remains a work-in-progress, we are excited to continue its development and further cultivate its use within the community.
The BLIS framework is primarily developed and maintained by individuals in the Science of High-Performance Computing (SHPC) group in the Oden Institute for Computational Engineering and Sciences at The University of Texas at Austin and in the Matthews Research Group at Southern Methodist University. Please visit the SHPC website for more information about our research group, such as a list of people and collaborators, funding sources, publications, and other educational projects (such as MOOCs).
Support for BLIS development at SMU was provided by the Office of Information Technology Research Technology Services team and HPC System Administrators on computational resources provided in partnership with SMU’s O’Donnell Data Science and Research Computing Institute.
Want to understand what's under the hood? Many of the same concepts and principles employed when developing BLIS are introduced and taught in a basic pedagogical setting as part of LAFF-On Programming for High Performance (LAFF-On-PfHP), one of several massive open online courses (MOOCs) in the Linear Algebra: Foundations to Frontiers series, all of which are available for free via the edX platform.
Plugin feature now available! BLIS addons (see below) provided a way to quickly extend BLIS's operation support or define new custom BLIS APIs for your application. BLIS plugins extend this support to completely external code, needing only an installed BLIS package (no source required). BLIS plugins also allow users to define their own kernels and blocksizes, combined with the cross-architecture support provided by the BLIS framework. Finally, user plugins can utilize the new API for modifying the BLIS "control tree" which defines the mathematical operation to be computed, as well as information controlling packing, partitioning, etc. Users can now modify the control tree to implement new linear algebra operations not already included in BLIS. See the documentation for an overview of these features and a step-by-step guides for creating plugins and modifying the control tree to implement an example operation "SYRKD".
BLIS selected for the 2023 James H. Wilkinson Prize for Numerical Software! We are thrilled to announce that Field Van Zee and Devin Matthews were chosen to receive the 2023 James H. Wilkinson Prize for Numerical Software. The selection committee sought to recognize the recipients "for the development of BLIS, a portable open-source software framework that facilitates rapid instantiation of high-performance BLAS and BLAS-like operations targeting modern CPUs." This prize is awarded once every four years to the authors of an outstanding piece of numerical software, or to individuals who have made an outstanding contribution to an existing piece of numerical software. It is awarded to an entry that best addresses all phases of the preparation of high-quality numerical software, and is intended to recognize innovative software in scientific computing and to encourage researchers in the earlier stages of their career. The prize will be awarded at the 2023 SIAM Conference on Computational Science and Engineering in Amsterdam.
Join us on Discord! In 2021, we soft-launched our Discord server by privately inviting current and former collaborators, attendees of our BLIS Retreat, as well as other participants within the BLIS ecosystem. We've been thrilled by the results thus far, and are happy to announce that our new community is now open to the broader public! If you'd like to hang out with other BLIS users and developers, ask a question, discuss future features, or just say hello, please feel free to join us! We've put together a step-by-step guide for creating an account and joining our cozy enclave. We even have a monthly "BLIS happy hour" event where people can casually come together for a video chat, Q&A, brainstorm session, or whatever it happens to unfold into!
Addons feature now available! Have you ever wanted to quickly extend BLIS's
operation support or define new custom BLIS APIs for your application, but were
unsure of how to add your source code to BLIS? Do you want to isolate your custom
code so that it only gets enabled when the user requests it? Do you like
sandboxes, but wish you didn't have to provide an
implementation of gemm? If so, you should check out our new
addons feature. Addons act like optional extensions that can be
created, enabled, and combined to suit your application's needs, all without
formally integrating your code into the core BLIS framework.
Multithreaded small/skinny matrix support for sgemm now available! Thanks to
funding and hardware support from Oracle, we have now accelerated gemm for
single-precision real matrix problems where one or two dimensions is exceedingly
small. This work is similar to the gemm optimization announced last year.
For now, we have only gathered performance results on an AMD Epyc Zen2 system, but
we hope to publish additional graphs for other architectures in the future. You may
find these Zen2 graphs via the PerformanceSmall document.
BLIS awarded SIAM Activity Group on Supercomputing Best Paper Prize for 2020! We are thrilled to announce that the paper that we internally refer to as the second BLIS paper,
"The BLIS Framework: Experiments in Portability." Field G. Van Zee, Tyler Smith, Bryan Marker, Tze Meng Low, Robert A. van de Geijn, Francisco Igual, Mikhail Smelyanskiy, Xianyi Zhang, Michael Kistler, Vernon Austel, John A. Gunnels, Lee Killough. ACM Transactions on Mathematical Software (TOMS), 42(2):12:1--12:19, 2016.
was selected for the SIAM Activity Group on Supercomputing Best Paper Prize for 2020. The prize is awarded once every two years to a paper judged to be the most outstanding paper in the field of parallel scientific and engineering computing, and has only been awarded once before (in 2016) since its inception in 2015 (the committee did not award the prize in 2018). The prize was awarded at the 2020 SIAM Conference on Parallel Processing for Scientific Computing in Seattle. Robert was present at the conference to give a talk on BLIS and accept the prize alongside other coauthors. The selection committee sought to recognize the paper, "which validates BLIS, a framework relying on the notion of microkernels that enables both productivity and high performance." Their statement continues, "The framework will continue having an important influence on the design and the instantiation of dense linear algebra libraries."
Multithreaded small/skinny matrix support for dgemm now available! Thanks to
contributions made possible by our partnership with AMD, we have dramatically
accelerated gemm for double-precision real matrix problems where one or two
dimensions is exceedingly small. A natural byproduct of this optimization is
that the traditional case of small m = n = k (i.e. square matrices) is also
accelerated, even though it was not targeted specifically. And though only
dgemm was optimized for now, support for other datatypes and/or other operations
may be implemented in the future. We've also added new graphs to the
PerformanceSmall document to showcase multithreaded
performance when one or more matrix dimensions are small.
Performance comparisons now available! We recently measured the per
暂无开放 Issues,或尚未同步最近议题。