We introduce Glimmer, a 10k base model trained on 500K tokens of FineWeb-Edu. The context window is 512 tokens The arch is standard llama (LlamaForCausalLM) 16 hidden dims 2 layers 4 attention heads 1 KV head (GQA) And the rest is on https://huggingface.co/Glint-Research/Glimmer-1-Base AMA for as lo