Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Papers Story
Papers

NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

Via ArXiv cs.CL
Monday, May 4, 2026 ยท 4:00AM
Summary

arXiv:2605.00086v1 Announce Type: new Abstract: High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder based on the ModernBERT architectu

Continue reading the full article
Read at ArXiv cs.CL
arxiv.org
Back to all stories