10 min read

Llama 3.1 and After: Open-Weight AI for WordPress Sites

Shashank Dubey
Content & Marketing, Wbcom Designs · Published Jul 26, 2024 · Updated Aug 29, 2026
WordPress Experts by Wbcom Designs - galaxy background with handwriting text

Llama 3.1 was the release that made open-weight language models a serious option for production work. When Meta published it in July 2024 it was the first time anyone could download a model (the 405B variant) that traded blows with GPT-4-class systems, and the smaller 8B and 70B versions became the default choice for teams that wanted to run a model on their own hardware. Two years on, the model itself has been superseded, but the pattern it established, open weights you can self-host or rent by the token, is exactly how most WordPress sites that use AI now do it.

This article covers what Llama 3.1 actually introduced, what Meta shipped after it (including the change of direction in 2026), and the practical ways a WordPress site owner can put a Llama-family model to work today.

What Llama 3.1 introduced

Card showing what Llama 3.1 introduced: 405B open weights, 128K context and eight languages.

Three things made the July 2024 release matter more than the Llama 2 and early Llama 3 releases before it.

  • A 405-billion-parameter model with open weights. Until then, frontier-scale models were API-only. Llama 3.1 405B could be downloaded from Hugging Face and run by anyone with enough GPUs, which in practice meant cloud providers and large companies, but it also meant hosted providers could compete on price for the same model.
  • 128K token context across all sizes. Llama 3 had shipped with 8K. Jumping to 128K on the 8B, 70B and 405B models meant you could feed a model a long document, a whole support thread, or a large chunk of a product catalogue in one request.
  • Official multilingual support and tool calling. Eight languages were supported out of the box (English, German, French, Italian, Portuguese, Hindi, Spanish and Thai), and the instruct models were trained to emit structured tool calls, which is what makes agent-style workflows possible.

The licence also loosened. The Llama 3.1 Community License allowed you to use model outputs to train other models, which Meta had previously forbidden, and the only hard restriction for commercial use was the 700 million monthly active user threshold that no WordPress site is going to hit.

For WordPress developers, the 8B instruct model was the interesting one. Quantised to 4-bit it fits in around 5GB, runs on a laptop with 16GB of RAM, and is good enough for summarising posts, classifying support tickets, drafting meta descriptions and moderating comments. That is the tier of task most sites actually need.

What came after: Llama 3.2, 3.3 and Llama 4

Timeline table from Llama 3.1 in July 2024 to Meta's Muse Glimmer 30B in August 2026.

Meta kept a fast cadence through 2025, then changed strategy in 2026. Here is the short version.

ReleaseDateWhat changedStill worth using?
Llama 3.1 (8B, 70B, 405B)July 2024128K context, 405B open weights, tool calling8B remains a fine small model; 70B and 405B are superseded
Llama 3.2 (1B, 3B, 11B Vision, 90B Vision)September 2024Tiny text models for edge devices; first Llama vision models1B and 3B are still popular for on-device and cheap classification
Llama 3.3 70BDecember 2024A 70B that matched 3.1 405B on most benchmarks at a fraction of the costYes. This is the workhorse many hosted providers still default to
Llama 4 Scout and MaverickApril 2025Mixture-of-experts (17B active parameters), native image input, 10M and 1M contextYes, via hosted APIs; too large to self-host casually
Llama 4 BehemothAnnounced 2025~2T parameters, never released publiclyNot available
Muse SparkApril 2026Meta Superintelligence Labs’ first model; closed weights, API and Meta apps onlyOnly through Meta’s own products
Muse Glimmer 30BAugust 202630B dense multimodal model under Apache 2.0, 128K context, distilled from Muse SparkYes. The best open-weight Meta model you can self-host right now

The Llama 4 launch in April 2025 was Meta’s move to mixture-of-experts architecture. Scout has 109B total parameters but only 17B active per token, which keeps inference cost down; Maverick is around 400B total with 128 experts. Both accept images as well as text. The EU multimodal restriction in the Llama 4 licence caught some people out: if your company is domiciled in the EU, the licence excludes the vision features.

Then 2026 brought the plot twist. Meta reorganised its AI effort under Meta Superintelligence Labs and in April shipped Muse Spark, a closed-weight reasoning model with no published weights and no architecture paper. For a while it looked like the open-weight era at Meta was over. In August 2026 the company released Muse Glimmer, a 30B dense model under a plain Apache 2.0 licence, ungated on Hugging Face, which is a more permissive licence than any Llama release ever had. As we write this, Meta has also said it intends to open the weights for Muse Spark 1.2, but they are not published yet.

So the practical advice in August 2026: for self-hosting, look at Muse Glimmer 30B or Llama 3.3 70B if you have the hardware, and Llama 3.1 8B or 3.2 3B if you do not. For hosted APIs, Llama 4 Maverick and Llama 3.3 70B are the cheap, capable options.

Three ways to run a Llama model behind a WordPress site

There is no single right answer. It depends on traffic, budget, and how sensitive your data is.

Option 1: Self-host with Ollama

Ollama is the easiest way to run open-weight models locally or on a VPS. It exposes an OpenAI-compatible API on port 11434, so any WordPress plugin that lets you set a custom OpenAI endpoint can talk to it. Setup on a Linux server:

# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Pull a model. llama3.1:8b is ~4.7GB, llama3.3:70b is ~43GB.
ollama pull llama3.1:8b

# Quick test
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"Write a 150 character meta description for a post about BuddyPress group settings."}]}'

Hardware is the catch. The 8B model runs acceptably on CPU with 16GB of RAM (a few tokens per second), comfortably on any GPU with 8GB of VRAM. The 70B needs a 48GB GPU or two 24GB cards. Muse Glimmer 30B at 4-bit fits in under 20GB, which puts it in reach of a single RTX 4090 or a Mac Studio. Do not run Ollama on the same box as WordPress unless it is a dedicated server with headroom; model inference will starve PHP of CPU.

Self-hosting makes sense when you process member data you would rather not send to a third party (private community content, LMS submissions, health or legal intake forms), or when volume is high enough that per-token pricing adds up.

Option 2: Rent the model from a hosted provider

If you do not want to manage GPUs, several providers serve Llama models by the token. Pricing as of spring 2026 per million tokens (input/output): Groq charges roughly $0.11/$0.34 for Llama 4 Scout and $0.50/$0.77 for Maverick; Together AI is $0.18/$0.59 and $0.27/$0.85; DeepInfra is usually the cheapest at around $0.08/$0.30 for Scout. AWS Bedrock, Azure AI Foundry and Google Vertex AI carry the same models at somewhat higher rates with enterprise contracts attached.

For a site generating, say, 2,000 meta descriptions and 500 comment moderation calls a month, that is a few dollars. The hosted route is the right default for most sites. OpenRouter is worth a look too: one API key, one OpenAI-compatible endpoint, and access to every Llama variant plus hundreds of other models, which makes it easy to swap models without touching WordPress.

Option 3: Fine-tune and host your own

Skip this unless you have a specific, repetitive task and a few thousand labelled examples. Fine-tuning Llama 3.1 8B with LoRA on a single GPU is well documented and cheap, but for most WordPress use cases a good system prompt plus retrieval from your own content gets you 90 percent of the way. We have done fine-tunes for clients classifying support tickets into 40 categories; we have never needed one for content generation.

Connecting a Llama model to WordPress

WordPress 7.0 (May 2026) shipped a core AI Client and an Abilities API, which means plugins can register AI providers and features in a standard way rather than each one bundling its own API layer. Provider plugins then plug into it. This is still young, but it is where things are going.

Today, the plugins we would actually recommend for Llama-family models:

  • AI Provider for Ollama (free, WordPress.org) registers your Ollama server with the core AI client, so any feature built on the Abilities API can use whatever model you have pulled.
  • AI Engine by Meow Apps is the most flexible general-purpose option. Under Meow Apps → AI Engine → Settings → Environments, add an environment with the OpenAI type, set the endpoint to http://your-server:11434/v1 (or your hosted provider’s URL) and add a model name. Chatbots, content assistant and the REST API then all route through it.
  • Ultimate AI Connector for Compatible Endpoints and Aiify both support Ollama and OpenRouter directly if you want something lighter.

If you are writing your own integration, keep it to a small mu-plugin that calls the endpoint with wp_remote_post(), caches results with transients, and never runs on a front-end page load. A typical pattern is to queue the job with Action Scheduler when a post is saved and write the result to post meta.

$response = wp_remote_post( 'http://10.0.0.5:11434/v1/chat/completions', array(
    'timeout' => 60,
    'headers' => array( 'Content-Type' => 'application/json' ),
    'body'    => wp_json_encode( array(
        'model'    => 'llama3.1:8b',
        'messages' => array(
            array( 'role' => 'system', 'content' => 'You write concise, plain meta descriptions under 155 characters.' ),
            array( 'role' => 'user',   'content' => wp_strip_all_tags( $post->post_content ) ),
        ),
    ) ),
) );

Sanitise the model’s output before you store or display it (wp_kses_post at minimum), and treat it like any other untrusted input. Prompt injection through user-submitted content is a real problem on community sites where members can write the text the model reads.

What this looks like on a community or LMS site

The uses that have held up in our projects, in rough order of value:

  1. Moderation triage. Run new activity posts and comments through a small model with a strict classification prompt. Flag, do not auto-delete. Llama 3.1 8B handles this well and the cost is negligible.
  2. Summaries of long threads and group discussions. The 128K context that 3.1 introduced is exactly what makes this workable; you can hand over an entire forum topic at once.
  3. Course assistant for LMS sites. Retrieval over lesson content plus a Llama 3.3 70B or Llama 4 model answering student questions, restricted to enrolled users. Self-hosting is attractive here because student data stays on your infrastructure.
  4. Editorial helpers. Tags, categories, excerpts and alt text suggestions in the editor. Cheap, boring, and saves real time.

We went into the moderation and chatbot side in more detail in what AI agents can do in a BuddyPress community. If you run a BuddyNext or BuddyX-based site, the same endpoint approach works because everything goes through standard WordPress hooks.

Frequently asked questions

Is Llama 3.1 still worth using?

The 8B model, yes, as a cheap local workhorse. The 70B and 405B have been overtaken by Llama 3.3 70B and Llama 4, which are cheaper to run for equal or better output. If you are starting fresh in 2026, begin with Llama 3.3 70B or Muse Glimmer 30B.

Is Llama “open source”?

Not by the OSI definition. The Llama licences are permissive community licences with usage restrictions and an attribution requirement (“Built with Llama”). Muse Glimmer, released under Apache 2.0, is genuinely open source. For a WordPress site the practical difference is small; for a product you redistribute, read the licence.

How much does self-hosting cost compared with the API?

A VPS with a 24GB GPU costs roughly $250 to $400 a month from providers like Hetzner, Lambda or RunPod. At Groq or DeepInfra prices that buys hundreds of millions of tokens, so the API is cheaper unless you have privacy requirements or very high volume.

Can I run a model on shared WordPress hosting?

No. Inference needs RAM and ideally a GPU that shared hosts do not provide. Put the model on a separate server (or use a hosted API) and have WordPress call it over HTTPS.

What we’d do

Four-step plan for a first Llama 3.1 style integration on WordPress, from hosted API to self-hosting.

Start with a hosted provider and an OpenAI-compatible plugin, pick Llama 3.3 70B or Llama 4 Maverick, and wire up one concrete task such as moderation flags or excerpt suggestions. Measure how often the output is useful. Only then decide whether self-hosting earns its keep, and if it does, Muse Glimmer 30B on a single-GPU box is the 2026 sweet spot. If you want help designing the integration, the custom plugin development team has built this stack for community and LMS sites.

Shashank Dubey
Content & Marketing, Wbcom Designs

Shashank Dubey, a contributor of Wbcom Designs is a blogger and a digital marketer. He writes articles associated with different niches such as WordPress, SEO, Marketing, CMS, Web Design, and Development, and many more.

Related reading