Skip to content

Small, specialised models

When a 6.75M-parameter transformer beats a frontier model: numeric transformers, column tokenisation, synthetic causal data, and training on consumer hardware. Grounded in MNM, which processes 43M+ prior authorisations a year.

Working on small, specialised models?