10 min read
Llama 3.1 and After: Open-Weight AI for WordPress Sites
Llama 3.1 was the release that made open-weight language models a serious option for production work. When Meta published it in July 2024 it was the first time anyone could download a model (the 405B variant) that traded blows with GPT-4-class systems, and the smaller 8B and 70B versions became the default choice for teams that wanted to run a model on their own hardware. Two years on, the model itself has been superseded, but the pattern it established, open weights you can self-host or rent by the token, is exactly how most WordPress sites that use AI now do it.
This article covers what Llama 3.1 actually introduced, what Meta shipped after it (including the change of direction in 2026), and the practical ways a WordPress site owner can put a Llama-family model to work today.
What Llama 3.1 introduced

Three things made the July 2024 release matter more than the Llama 2 and early Llama 3 releases before it.
- A 405-billion-parameter model with open weights. Until then, frontier-scale models were API-only. Llama 3.1 405B could be downloaded from Hugging Face and run by anyone with enough GPUs, which in practice meant cloud providers and large companies, but it also meant hosted providers could compete on price for the same model.
- 128K token context across all sizes. Llama 3 had shipped with 8K. Jumping to 128K on the 8B, 70B and 405B models meant you could feed a model a long document, a whole support thread, or a large chunk of a product catalogue in one request.
- Official multilingual support and tool calling. Eight languages were supported out of the box (English, German, French, Italian, Portuguese, Hindi, Spanish and Thai), and the instruct models were trained to emit structured tool calls, which is what makes agent-style workflows possible.
The licence also loosened. The Llama 3.1 Community License allowed you to use model outputs to train other models, which Meta had previously forbidden, and the only hard restriction for commercial use was the 700 million monthly active user threshold that no WordPress site is going to hit.
For WordPress developers, the 8B instruct model was the interesting one. Quantised to 4-bit it fits in around 5GB, runs on a laptop with 16GB of RAM, and is good enough for summarising posts, classifying support tickets, drafting meta descriptions and moderating comments. That is the tier of task most sites actually need.
What came after: Llama 3.2, 3.3 and Llama 4

Meta kept a fast cadence through 2025, then changed strategy in 2026. Here is the short version.
| Release | Date | What changed | Still worth using? |
|---|---|---|---|
| Llama 3.1 (8B, 70B, 405B) | July 2024 | 128K context, 405B open weights, tool calling | 8B remains a fine small model; 70B and 405B are superseded |
| Llama 3.2 (1B, 3B, 11B Vision, 90B Vision) | September 2024 | Tiny text models for edge devices; first Llama vision models | 1B and 3B are still popular for on-device and cheap classification |
| Llama 3.3 70B | December 2024 | A 70B that matched 3.1 405B on most benchmarks at a fraction of the cost | Yes. This is the workhorse many hosted providers still default to |
| Llama 4 Scout and Maverick | April 2025 | Mixture-of-experts (17B active parameters), native image input, 10M and 1M context | Yes, via hosted APIs; too large to self-host casually |
| Llama 4 Behemoth | Announced 2025 | ~2T parameters, never released publicly | Not available |
| Muse Spark | April 2026 | Meta Superintelligence Labs’ first model; closed weights, API and Meta apps only | Only through Meta’s own products |
| Muse Glimmer 30B | August 2026 | 30B dense multimodal model under Apache 2.0, 128K context, distilled from Muse Spark | Yes. The best open-weight Meta model you can self-host right now |
The Llama 4 launch in April 2025 was Meta’s move to mixture-of-experts architecture. Scout has 109B total parameters but only 17B active per token, which keeps inference cost down; Maverick is around 400B total with 128 experts. Both accept images as well as text. The EU multimodal restriction in the Llama 4 licence caught some people out: if your company is domiciled in the EU, the licence excludes the vision features.
Then 2026 brought the plot twist. Meta reorganised its AI effort under Meta Superintelligence Labs and in April shipped Muse Spark, a closed-weight reasoning model with no published weights and no architecture paper. For a while it looked like the open-weight era at Meta was over. In August 2026 the company released Muse Glimmer, a 30B dense model under a plain Apache 2.0 licence, ungated on Hugging Face, which is a more permissive licence than any Llama release ever had. As we write this, Meta has also said it intends to open the weights for Muse Spark 1.2, but they are not published yet.
So the practical advice in August 2026: for self-hosting, look at Muse Glimmer 30B or Llama 3.3 70B if you have the hardware, and Llama 3.1 8B or 3.2 3B if you do not. For hosted APIs, Llama 4 Maverick and Llama 3.3 70B are the cheap, capable options.
Three ways to run a Llama model behind a WordPress site
There is no single right answer. It depends on traffic, budget, and how sensitive your data is.
Option 1: Self-host with Ollama
Ollama is the easiest way to run open-weight models locally or on a VPS. It exposes an OpenAI-compatible API on port 11434, so any WordPress plugin that lets you set a custom OpenAI endpoint can talk to it. Setup on a Linux server:
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Pull a model. llama3.1:8b is ~4.7GB, llama3.3:70b is ~43GB.
ollama pull llama3.1:8b
# Quick test
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"Write a 150 character meta description for a post about BuddyPress group settings."}]}'
Hardware is the catch. The 8B model runs acceptably on CPU with 16GB of RAM (a few tokens per second), comfortably on any GPU with 8GB of VRAM. The 70B needs a 48GB GPU or two 24GB cards. Muse Glimmer 30B at 4-bit fits in under 20GB, which puts it in reach of a single RTX 4090 or a Mac Studio. Do not run Ollama on the same box as WordPress unless it is a dedicated server with headroom; model inference will starve PHP of CPU.
Self-hosting makes sense when you process member data you would rather not send to a third party (private community content, LMS submissions, health or legal intake forms), or when volume is high enough that per-token pricing adds up.
Option 2: Rent the model from a hosted provider
If you do not want to manage GPUs, several providers serve Llama models by the token. Pricing as of spring 2026 per million tokens (input/output): Groq charges roughly $0.11/$0.34 for Llama 4 Scout and $0.50/$0.77 for Maverick; Together AI is $0.18/$0.59 and $0.27/$0.85; DeepInfra is usually the cheapest at around $0.08/$0.30 for Scout. AWS Bedrock, Azure AI Foundry and Google Vertex AI carry the same models at somewhat higher rates with enterprise contracts attached.
For a site generating, say, 2,000 meta descriptions and 500 comment moderation calls a month, that is a few dollars. The hosted route is the right default for most sites. OpenRouter is worth a look too: one API key, one OpenAI-compatible endpoint, and access to every Llama variant plus hundreds of other models, which makes it easy to swap models without touching WordPress.
Option 3: Fine-tune and host your own
Skip this unless you have a specific, repetitive task and a few thousand labelled examples. Fine-tuning Llama 3.1 8B with LoRA on a single GPU is well documented and cheap, but for most WordPress use cases a good system prompt plus retrieval from your own content gets you 90 percent of the way. We have done fine-tunes for clients classifying support tickets into 40 categories; we have never needed one for content generation.
Connecting a Llama model to WordPress
WordPress 7.0 (May 2026) shipped a core AI Client and an Abilities API, which means plugins can register AI providers and features in a standard way rather than each one bundling its own API layer. Provider plugins then plug into it. This is still young, but it is where things are going.
Today, the plugins we would actually recommend for Llama-family models:
- AI Provider for Ollama (free, WordPress.org) registers your Ollama server with the core AI client, so any feature built on the Abilities API can use whatever model you have pulled.
- AI Engine by Meow Apps is the most flexible general-purpose option. Under Meow Apps → AI Engine → Settings → Environments, add an environment with the OpenAI type, set the endpoint to
http://your-server:11434/v1(or your hosted provider’s URL) and add a model name. Chatbots, content assistant and the REST API then all route through it. - Ultimate AI Connector for Compatible Endpoints and Aiify both support Ollama and OpenRouter directly if you want something lighter.
If you are writing your own integration, keep it to a small mu-plugin that calls the endpoint with wp_remote_post(), caches results with transients, and never runs on a front-end page load. A typical pattern is to queue the job with Action Scheduler when a post is saved and write the result to post meta.
$response = wp_remote_post( 'http://10.0.0.5:11434/v1/chat/completions', array(
'timeout' => 60,
'headers' => array( 'Content-Type' => 'application/json' ),
'body' => wp_json_encode( array(
'model' => 'llama3.1:8b',
'messages' => array(
array( 'role' => 'system', 'content' => 'You write concise, plain meta descriptions under 155 characters.' ),
array( 'role' => 'user', 'content' => wp_strip_all_tags( $post->post_content ) ),
),
) ),
) );
Sanitise the model’s output before you store or display it (wp_kses_post at minimum), and treat it like any other untrusted input. Prompt injection through user-submitted content is a real problem on community sites where members can write the text the model reads.
What this looks like on a community or LMS site
The uses that have held up in our projects, in rough order of value:
- Moderation triage. Run new activity posts and comments through a small model with a strict classification prompt. Flag, do not auto-delete. Llama 3.1 8B handles this well and the cost is negligible.
- Summaries of long threads and group discussions. The 128K context that 3.1 introduced is exactly what makes this workable; you can hand over an entire forum topic at once.
- Course assistant for LMS sites. Retrieval over lesson content plus a Llama 3.3 70B or Llama 4 model answering student questions, restricted to enrolled users. Self-hosting is attractive here because student data stays on your infrastructure.
- Editorial helpers. Tags, categories, excerpts and alt text suggestions in the editor. Cheap, boring, and saves real time.
We went into the moderation and chatbot side in more detail in what AI agents can do in a BuddyPress community. If you run a BuddyNext or BuddyX-based site, the same endpoint approach works because everything goes through standard WordPress hooks.
Frequently asked questions
Is Llama 3.1 still worth using?
The 8B model, yes, as a cheap local workhorse. The 70B and 405B have been overtaken by Llama 3.3 70B and Llama 4, which are cheaper to run for equal or better output. If you are starting fresh in 2026, begin with Llama 3.3 70B or Muse Glimmer 30B.
Is Llama “open source”?
Not by the OSI definition. The Llama licences are permissive community licences with usage restrictions and an attribution requirement (“Built with Llama”). Muse Glimmer, released under Apache 2.0, is genuinely open source. For a WordPress site the practical difference is small; for a product you redistribute, read the licence.
How much does self-hosting cost compared with the API?
A VPS with a 24GB GPU costs roughly $250 to $400 a month from providers like Hetzner, Lambda or RunPod. At Groq or DeepInfra prices that buys hundreds of millions of tokens, so the API is cheaper unless you have privacy requirements or very high volume.
Can I run a model on shared WordPress hosting?
No. Inference needs RAM and ideally a GPU that shared hosts do not provide. Put the model on a separate server (or use a hosted API) and have WordPress call it over HTTPS.
What we’d do

Start with a hosted provider and an OpenAI-compatible plugin, pick Llama 3.3 70B or Llama 4 Maverick, and wire up one concrete task such as moderation flags or excerpt suggestions. Measure how often the output is useful. Only then decide whether self-hosting earns its keep, and if it does, Muse Glimmer 30B on a single-GPU box is the 2026 sweet spot. If you want help designing the integration, the custom plugin development team has built this stack for community and LMS sites.
Related reading