Research Seminar

The IOL Seminar and Lecture Series at the Zuse Institute Berlin serves to bring together researchers presenting their latest work and to organize tutorial lectures on valuable topics not typically covered in graduate coursework. Presentations usually take place on Wednesday afternoons in ZIB's Seminar Room 2006.

For talks before year 2026, see 2024–2025, 2022–2023, and 2019–2021.


Photo Moritz Wagner Moritz Wagner – ZIB
@ ZIB, Room 4027 (Roter Salon)

Title. A Free Lunch in LLM Compression: Revisiting Retraining after Pruning

Abstract.

Post-training pruning removes weights from a trained language model to cut inference cost, but the pruned model loses quality unless the remaining weights are adapted. For decades the standard answer was to prune and then retrain. However, for large language models (LLMs), retraining was declared infeasible, and the field responded with developing ever more elaborate rules for choosing which weights to remove so that no adaptation is needed afterwards. We argue that post-pruning adaptation is still feasible in the era of LLMs. We revisit local reconstruction: after pruning with a fixed mask, one submodel at a time is adapted on a few hundred calibration sequences to match the intermediate activations of the dense model. Across four model families from 0.5B to 72B parameters, we establish three findings. First, local reconstruction matches LoRA-style retraining in perplexity and zero-shot accuracy while using 64 times fewer samples and up to 130 times less compute, and it fits a 32B model on a single 80 GB GPU. Second, the size of the reconstructed submodel, from half a transformer block to a quarter of the network, has almost no effect on final quality, while peak memory varies by hundreds of gigabytes. The one exception is per-matrix reconstruction, the formulation most widely used in the literature, which consistently underperforms. We trace this failure to compositional error accumulation and show that including a single nonlinearity in the reconstructed submodel is what removes it. Third, once reconstruction is applied, the gap between sophisticated pruning criteria and plain magnitude pruning shrinks with model scale and essentially vanishes above 30B parameters. Part of what sophisticated criteria bought was compensation for a missing adaptation step. Together, these results establish local reconstruction as a practical default for post-pruning adaptation at LLM scale and shift the central question of LLM pruning from which weights to remove to how to adapt the ones that remain.


Photo Ingo Meise Ingo Meise – ZIB
@ ZIB, Room 4027 (Roter Salon)

Title. Tight PAC Guarantees for Smooth Boosting and a New Smoothness Hyperparameter

Abstract.

The classical weak learning setting has had a considerable impact on both statistical learning theory and the development of ensemble learning algorithms. Given a binary classification dataset and an input distribution over the data points, the weak learner outputs a classifier with an edge over random guessing. However, this edge can become too small to ensure fast convergence and good generalization. To address this risk, several methods have been developed that limit the influence of outliers, thereby achieving faster convergence. A proof of improved generalization, however, has only been provided under the assumption of label noise. Against this background, we propose a one-dimensional smoothness parameter and a corresponding smooth boosting routine, and we prove a generalization guarantee that is tight for all values of this parameter. Moreover, this guarantee can be remarkably better than the recently established tight weak-to-strong learning guarantee for the classical non-smooth weak learning setting, which is recovered by setting the smoothness parameter to zero. We then go a step further and apply this smoothness parameter in the context of fastest boosting, in which the computed ensemble can use arbitrary aggregation rules and is not restricted to weighted combinations. The practical usefulness of the proposed smoothness parameter is validated in extensive experiments.


Photo Karl Welzel Karl Welzel – University of Oxford
@ ZIB, Room 4027 (Roter Salon)

Title. Do third derivatives accelerate unconstrained continuous optimization?

Abstract.

Methods for unconstrained continuous optimization can be classified by the strength of the oracle for the objective function they require. First-order methods only require access to function values and gradients, second-order methods additionally require access to Hessians. Higher-order methods correspondingly require access to derivatives of order 1 to p for some p ≥ 3 and are called tensors methods. Intuitively, access to additional derivative information should speed up convergence to an approximate minimizer. We will cover the different perspectives from which this becomes true. First, convergence speed can be quantified from the perspective of global iteration complexity, for which the literature asserts that methods minimizing a regularized Taylor expansions, and in particular the adaptive ARp method, are optimal. Second, we will discuss local convergence results, showing how ARp can converge superlinearly even if the Hessian is singular at the minimizer. Third, numerical experiments on a range of test functions show that (after introducing certain heuristics) the third-order AR3 method needs fewer iterations and oracle calls than the second-order AR2 method.


Photo Rebekka Burkholz Rebekka Burkholz – CISPA Helmholtz Center for Information Security
@ ZIB, Room 0001 (Studio da Vinci)

Title. Towards AI That Is Smart, Sparse, and Social

Abstract.

Deep learning continues to achieve impressive breakthroughs across disciplines but relies on increasingly large neural network models that are trained on massive data sets. Their development inflicts costs that are only affordable by a few labs and prevent global participation in the creation of related technologies. But does it really have to be like this? We will identify some of the major challenges of deep learning at small scales and present solution strategies pertaining to the design of sparse training algorithms and problem specific neural network design in the context of agentic networks, which hold the promise to overcome a fundamental trade-off between model specialization and trainability.