Search Results for author: Sundar Subramanian

Found 1 papers, 0 papers with code

FP8 versus INT8 for efficient deep learning inference

no code implementations • 31 Mar 2023 • Mart van Baalen, Andrey Kuzmin, Suparna S Nair, Yuwei Ren, Eric Mahurin, Chirag Patel, Sundar Subramanian, Sanghyuk Lee, Markus Nagel, Joseph Soriaga, Tijmen Blankevoort

We theoretically show the difference between the INT and FP formats for neural networks and present a plethora of post-training quantization and quantization-aware-training results to show how this theory translates to practice.

Quantization

Paper
Add Code

Cannot find the paper you are looking for? You can Submit a new open access paper.