Search Results for author: Satrajit Chatterjee

Found 11 papers, 3 papers with code

A Closer Look at Hardware-Friendly Weight Quantization

1 code implementation • 7 Oct 2022 • Sungmin Bae, Piotr Zielinski, Satrajit Chatterjee

We study the two methods on MobileNetV1 and MobileNetV2 using multiple empirical metrics to identify the sources of performance differences between the two classes, namely, sensitivity to outliers and convergence instability of the quantizer scaling factor.

Quantization

521

Paper
Code

On the Generalization Mystery in Deep Learning

no code implementations • 18 Mar 2022 • Satrajit Chatterjee, Piotr Zielinski

The generalization mystery in deep learning is the following: Why do over-parameterized neural networks trained with gradient descent (GD) generalize well on real datasets even though they are capable of fitting random datasets of comparable size?

Memorization

Paper
Add Code

Enabling Binary Neural Network Training on the Edge

2 code implementations • 8 Feb 2021 • Erwei Wang, James J. Davis, Daniele Moro, Piotr Zielinski, Jia Jie Lim, Claudionor Coelho, Satrajit Chatterjee, Peter Y. K. Cheung, George A. Constantinides

The ever-growing computational demands of increasingly complex machine learning models frequently necessitate the use of powerful cloud-based infrastructure for their training.

Quantization

521

Paper
Code

Apollo: Transferable Architecture Exploration

no code implementations • 2 Feb 2021 • Amir Yazdanbakhsh, Christof Angermueller, Berkin Akin, Yanqi Zhou, Albin Jones, Milad Hashemi, Kevin Swersky, Satrajit Chatterjee, Ravi Narayanaswami, James Laudon

We further show that by transferring knowledge between target architectures with different design constraints, Apollo is able to find optimal configurations faster and often with better objective value (up to 25% improvements).

Paper
Add Code

Making Coherence Out of Nothing At All: Measuring Evolution of Gradient Alignment

no code implementations • 1 Jan 2021 • Satrajit Chatterjee, Piotr Zielinski

Using $m$-coherence, we study the evolution of alignment of per-example gradients in ResNet and EfficientNet models on ImageNet and several variants with label noise, particularly from the perspective of the recently proposed Coherent Gradients (CG) theory that provides a simple, unified explanation for memorization and generalization [Chatterjee, ICLR 20].

Memorization

Paper
Add Code

Logic Synthesis Meets Machine Learning: Trading Exactness for Generalization

1 code implementation • 4 Dec 2020 • Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi Masahiro Fujita, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Junior, Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong R. Jiang, Jiaqi Gu, Zheng Zhao, Zixuan Jiang, David Z. Pan, Brunno A. de Abreu, Isac de Souza Campos, Augusto Berndt, Cristina Meinhardt, Jonata T. Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar, Wei Zeng, Azadeh Davoodi, Rasit O. Topaloglu, Yuan Zhou, Jordan Dotzel, Yichi Zhang, Hanyu Wang, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee

If the function is incompletely-specified, the implementation has to be true only on the care set.

BIG-bench Machine Learning

Paper
Code

Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment

no code implementations • 3 Aug 2020 • Satrajit Chatterjee, Piotr Zielinski

Using $m$-coherence, we study the evolution of alignment of per-example gradients in ResNet and Inception models on ImageNet and several variants with label noise, particularly from the perspective of the recently proposed Coherent Gradients (CG) theory that provides a simple, unified explanation for memorization and generalization [Chatterjee, ICLR 20].

Memorization

Paper
Add Code

Weak and Strong Gradient Directions: Explaining Memorization, Generalization, and Hardness of Examples at Scale

no code implementations • 16 Mar 2020 • Piotr Zielinski, Shankar Krishnan, Satrajit Chatterjee

The key insight of CGH is that, since the overall gradient for a single step of SGD is the sum of the per-example gradients, it is strongest in directions that reduce the loss on multiple examples if such directions exist.

Memorization

Paper
Add Code

Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization

no code implementations • ICLR 2020 • Satrajit Chatterjee

We propose an approach to answering this question based on a hypothesis about the dynamics of gradient descent that we call Coherent Gradients: Gradients from similar examples are similar and so the overall gradient is stronger in certain directions where these reinforce each other.

Descriptive Open-Ended Question Answering

Paper
Add Code

Circuit-Based Intrinsic Methods to Detect Overfitting

no code implementations • ICML 2020 • Satrajit Chatterjee, Alan Mishchenko

By intrinsic methods, we mean methods that rely only on the model and the training data, as opposed to traditional methods (we call them extrinsic methods) that rely on performance on a test set or on bounds from model complexity.

counterfactual Memorization

Paper
Add Code

Learning and Memorization

no code implementations • ICML 2018 • Satrajit Chatterjee

In the machine learning research community, it is generally believed that there is a tension between memorization and generalization.

Memorization

Paper
Add Code

Cannot find the paper you are looking for? You can Submit a new open access paper.