Research output

Publications

Selected papers, preprints, and conference contributions related to the project.

Showing 12 of 13 publications
Publication2026

Amortizing Scaling Law Construction Costs

AK Jha, DA Onuţu, N Mallik, S Haldar, S Laing…

Scaling laws guide the design choices for training large foundation models, but deriving them involves training an exhaustive grid over hyperparameters, token budgets, and parameter …

Publication2026

Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

T Mickus, C Savelli, E Calò, E Raimond, S Frank…

In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-…

Publication2026

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

R Sourty, A Chaffin, PRM Junior…

State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study …

Publication2026

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

N Ajroldi, DA Onutu, H Al-Tahan, J Franke…

We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling jointly optimal learning rates …

Publication2026

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

S Riazhskykh, N Luu, O Bojar

Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, ie, save up to 33% of training tokens needed to reach a …

Publication2026

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

E Lushtaku, B Kargi, A Elganzory, F Ferreira…

LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, …

Publication2026

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

A Hochlehnert, M Nezhurina, M Cherti…

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we …

Publication2026

Opensubtitles2024: A massively parallel dataset of movie subtitles for mt development and evaluation

J Tiedemann, H Luo

This paper introduces OpenSubtitles2024, a massively parallel dataset compiled from translated subtitles. The collection includes an extensive collection of aligned training data based …

Publication2026

Openthoughts: Data recipes for reasoning models

E Guha, R Marten, S Keh, N Raoof…

Abstract Reasoning models have made rapid progress on many benchmarks involving math, code, and science. Yet, there are still many open questions about the best train-ing recipes …

Publication2026

Training dynamics impact post-training quantization robustness

A Catalan-Tatjer, N Ajroldi…

While post-training quantization is widely adopted for efficient deployment of large language models, the mechanisms underlying quantization robustness remain unclear. We conduct a …

Publication2025

FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models

J KytÃķniemi, J Piha, A Reunamo, F Vitiugin…

We introduce FIN-bench-v2, a unified benchmark suite for evaluating large language models in Finnish. FIN-bench-v2 consolidates Finnish versions of widely used benchmarks …

Publication

MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models

S Ji, Z Li, J Paavola, P Lin, P Chen, D O'Brien…

In this work, we present the MaLA (Massively multilingual Language Adaptation) suite, a comprehensive collection of open-access resources designed to advance language technology …