Databricks releases DBRX
Built by the former MosaicML team Databricks had acquired the previous year, the 132-billion-parameter model activated only 36 billion parameters per input.
- Open weights & ecosystem
- Models & capabilities
- Minor
Databricks released DBRX, an open-weight, 132-billion-parameter mixture-of-experts language model, and said it outperformed Llama 2 70B, Mixtral and xAI’s Grok-1 on a set of standard benchmarks. The model was built by the team from MosaicML, the AI training-infrastructure company Databricks had acquired for $1.3 billion in 2023, and represented the acquisition’s first major model release under the Databricks name.
DBRX used a fine-grained mixture-of-experts architecture with 16 experts, of which 4 were active for any given input, so only about 36 billion of its 132 billion parameters were used per token — an efficiency approach similar in spirit to Mixtral’s but with more, smaller experts. Databricks said this design gave the model roughly twice the inference speed of a comparably capable dense model.
The release positioned Databricks, a data and analytics infrastructure company rather than a dedicated AI lab, as a direct entrant in the open-weight large-model race then being run by Meta, Mistral and xAI. Databricks pitched DBRX primarily at enterprise customers who wanted to fine-tune and deploy a capable open model on their own data within the Databricks platform, rather than at consumer or developer audiences chasing leaderboard position for its own sake — continuing the argument that open weights were becoming a viable commercial foundation, not just a research artefact.