N 41.053° · W 73.539°
/AI THOUGHT LEADERSHIP

Llama 4 Scout and Maverick: Meta's Open-Source Contenders

By Scott McKenna, Founder · 2026-04-12 · AI Thought Leadership · Updated May 13, 2026

What Meta actually shipped

Llama 4 arrived as a family rather than a single model, with two members released as downloadable weights: Scout and Maverick. Both are built on a mixture-of-experts design and both are natively multimodal, meaning they take images as well as text as input rather than having vision bolted on afterwards.

Mixture-of-experts is worth thirty seconds of explanation, because it is the reason the specifications look strange. Instead of every part of the model working on every word, the model is divided into many smaller "experts", and a router picks a handful for each token. The result is a model with a very large total parameter count but a much smaller number of parameters active at any moment. You get the knowledge of a big model at closer to the running cost of a small one, at the price of needing enough memory to hold the whole thing regardless.

Scout and Maverick, side by side

Scout

The smaller of the two, with a modest number of experts and a total size that Meta positioned as fitting on a single high-end data centre GPU once quantised. Its headline feature at launch was a very large advertised context window, which in principle allows enormous documents to be processed in one pass. Independent testing since has suggested that usable performance across the full advertised window is considerably less impressive than the number implies, which is a pattern seen across the industry rather than something unique to Meta.

Maverick

The larger sibling, with many more experts and a much larger total size, aimed at general assistant work and coding. Meta positioned it against the mid-tier commercial models rather than the absolute frontier. It performs respectably on general reasoning and coding tasks while remaining cheaper to serve than a dense model of comparable knowledge.

There was also a preview of a much larger model, Behemoth, discussed as a teacher model for the family. Treat anything about it as unconfirmed until weights actually appear.

The word "open source" needs an asterisk

This matters more than the benchmark charts for anyone considering these commercially.

Llama 4 is released under Meta's own community licence, not a standard open source licence such as Apache or MIT. The practical conditions include an attribution requirement, naming rules for derived models, a monthly-active-user threshold above which you must negotiate a separate licence with Meta, and additional geographic restrictions that have affected multimodal use for organisations domiciled in the European Union.

For a local business none of those thresholds will ever bind you. But "open weights" and "open source" are genuinely different things, and if you build a product on top of these models it is worth having someone read the licence properly rather than assuming it behaves like the open source software you are used to.

One further note on trust: at launch there was controversy over an experimental variant appearing on a public preference leaderboard that differed from the released weights. Meta acknowledged the variant was a different build. It is a reasonable reminder to test on your own tasks rather than trusting launch-day rankings, whoever publishes them.

What running one of these actually involves

The phrase "free model" does a lot of concealing. The weights cost nothing. Everything else does.

If you are weighing whether to buy hardware for this at all, our breakdown of AI costs for small businesses covers the comparison against simply paying per use.

Should a local business care?

Mostly, no, and it is worth saying plainly. If you use AI to draft emails, reply to reviews and write service pages, a commercial assistant subscription will serve you better for less effort, and the quality gap on those tasks is unlikely to be your bottleneck.

Where open-weight models genuinely earn their place is narrower: when data cannot leave your premises for regulatory or contractual reasons, when your usage volume is high and predictable enough that per-token pricing hurts, or when you need a model that will not change underneath you because a vendor shipped an update. Those are real situations, they are just not most situations.

The broader significance of the Llama family is competitive rather than practical. A capable model anyone can download puts a ceiling on what the commercial labs can charge, and that benefits you even if you never download anything.

Is Llama 4 really free to use?

The weights are free to download and Meta's licence permits commercial use for the overwhelming majority of businesses. It is not a standard open source licence, though: attribution and naming conditions apply, very large platforms must negotiate separately, and some geographic restrictions exist. Running it is not free either, since you pay for hardware or hosting.

What is the difference between Scout and Maverick?

Scout is the smaller model with fewer experts, designed to run on more modest hardware, and it launched with a very large advertised context window. Maverick is substantially larger with many more experts and generally performs better on reasoning and coding. Both share the same mixture-of-experts architecture and both accept images as well as text.

Should my small business run Llama 4 instead of paying for ChatGPT?

Usually not. A commercial subscription is cheaper once you count hardware, setup and maintenance, and it requires no technical upkeep. Open weights make sense when data genuinely cannot leave your premises, when volume is high and predictable, or when you need a model that will not change without warning.

Can I trust the benchmark scores Meta published?

Treat launch-day figures from any lab as marketing rather than measurement. In Llama 4's case there was specific controversy about an experimental variant appearing on a public leaderboard that differed from the released weights. The reliable approach is to run five of your own real tasks through the model and judge the output yourself.

Want this handled for you?

Get a free audit of your website, Google reviews, and local SEO — we’ll show you exactly where you’re losing customers. Delivered in 24 hours, no sales call.

Get my free audit → or book a 15-min call

Want AI Working for Your Business?

We help local businesses in Stamford, Greenwich, Norwalk, and Fairfield County implement AI marketing that generates real results.

Get Your Free AI Marketing Audit →
SERVICES: Digital Marketing SEO Services Google Ads LOCATIONS: Stamford Greenwich Norwalk White Plains RESOURCES: Blog Free Audit Free Tools