Publications
Empirical Inference
Deep Models and Optimization
Conference Paper
Scaling Behavior of Discrete Diffusion Language Models
von Rütte, D., Fluri, J., Pooladzandi, O., Schölkopf, B., Hofmann, T., Orvieto, A.
The Fourteenth International Conference on Learning Representations (ICLR), April 2026 (Published)
arXiv
URL
BibTeX
Empirical Inference
Deep Models and Optimization
Conference Paper
Generalized Interpolating Discrete Diffusion
von Rütte, D., Fluri, J., Ding, Y., Orvieto, A., Schölkopf, B., Hofmann, T.
Proceedings of the 42nd International Conference on Machine Learning (ICML), 267:61810-61843, Proceedings of Machine Learning Research, (Editors: Singh, Aarti and Fazel, Maryam and Hsu, Daniel and Lacoste-Julien, Simon and Berkenkamp, Felix and Maharaj, Tegan and Wagstaff, Kiri and Zhu, Jerry), PMLR, International Conference on Machine Learning, July 2025 (Published)
arXiv
URL
BibTeX
Deep Models and Optimization
Conference Paper
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
Movahedi, S., Orvieto, A., Moosavi-Dezfooli, S.
In The Thirteenth International Conference on Learning Representations, ICLR 2025, The Thirteenth International Conference on Learning Representations, January 2025 (Accepted)
BibTeX
Deep Models and Optimization
Conference Paper
Using Shapley interactions to understand how models use structure
Divyansh Singhvi, D. M. A. E. R. J. I. P. N. S.
In Proceedings ACL, 1-20, Vienna Center, Association for Computational Linguistics (ACL 2025), 2025 (Accepted)
DOI
URL
BibTeX
Deep Models and Optimization
Article
An accelerated lyapunov function for Polyak’s Heavy-ball on convex quadratics
Orvieto, A.
Optimization Letters, 19:307-328, 2025 (Published)
DOI
URL
BibTeX
Deep Models and Optimization
Conference Paper
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
Monzio Compagnoni, E., Liu, T., Islamov, R., Proske, F. N., Orvieto, A., Lucchi, A.
In The Thirteenth International Conference on Learning Representations, ICLR 2025, The Thirteenth International Conference on Learning Representations, November 2024 (Accepted)
BibTeX
Deep Models and Optimization
Article
NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
Köprücü, N., Okpekpe, D., Orvieto, A.
October 2024 (In preparation)
BibTeX
Deep Models and Optimization
Conference Paper
Loss Landscape Characterization of Neural Networks without Over-Parametrization
Islamov, R., Ajroldi, N., Orvieto, A., Lucchi, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
Zucchet, N., Orvieto, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Theoretical Foundations of Deep Selective State-Space Models
Muca Cirone, N., Orvieto, A., Walker, B., Salvi, C., Lyons, T.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
Sieber, J., Amo Alonso, C., Didier, A., Zeilinger, M., Orvieto, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Article
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
Orvieto, A., Xiao, L.
July 2024 (Submitted)
BibTeX
Deep Models and Optimization
Article
An accelerated Lyapunov function for Polyak’s Heavy-ball on convex quadratics
Orvieto, A.
Optimization Letters, June 2024 (Published)
BibTeX
Deep Models and Optimization
Article
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
Yi Meng, S., Orvieto, A., Yiming Cao, D., De Sa, C.
June 2024 (Submitted)
BibTeX
Deep Models and Optimization
Conference Paper
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
Orvieto, A., De, S., Gulcehre, C., Pascanu, R., Smith, S. L.
In Proceedings of Machine Learning Research, Proceedings of the Forty-First International Conference on Machine Learning , Forty-First International Conference on Machine Learning , June 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
SDEs for Minimax Optimization
Monzio Compagnoni, E., Orvieto, A., Kersting, H., Proske, F., Lucchi, A.
PMLR, AISTATS, February 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Recurrent Distance Filtering for Graph Representation Learning
Ding, Y., Orvieto, A., He, B., Hofmann, T.
In PMLR, ICML, January 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Super Consistency of Neural Network Landscapes and Learning Rate Transfer
Noci, L., Meterez, A., Hofmann, T., Orvieto, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, January 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Resurrecting Recurrent Neural Networks for Long Sequences
Orvieto, A., Smith, S. L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R., De, S.
In Proceedings of the Eleventh International Conference on Learning Representations, ICLR, June 2023 (Published)
URL
BibTeX