AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Google announced EmbeddingGemma 2 on Oct. 6, 2026, an open-weight model that maps text, code, images, audio and video into a shared embedding space. The company says its 740-million-parameter model is designed for on-device use; benchmark and performance claims are Google’s, and independent results are not included in the announcement.

Google announced EmbeddingGemma 2 on Oct. 6, introducing a 740-million-parameter model that converts text, code, images, audio and video into a shared embedding space for search and retrieval. The model is released under the Apache 2.0 license and is designed to run on consumer devices, potentially letting developers build cross-media search tools without sending source content to a remote service.

Google says the model is built on the Gemma 4 architecture and extends EmbeddingGemma, which the company introduced the previous year for text embeddings. Google reports that the original model has been downloaded more than 20 million times; that figure is the company’s count, and the announcement does not provide a breakdown of usage or deployments.

EmbeddingGemma 2 can represent different types of material in one vector space. In practice, Google says a developer could build a tool that finds a video clip from a voice memo or searches audio recordings with a text query. The company describes an 8,192-token context window, which it says can cover up to 5.5 minutes of audio, 29 images, 58 video frames, or combinations of those inputs. These are stated capacity limits, not guarantees of search accuracy for every workload.

The announcement describes a modular design: text-only use can require as few as 270 million parameters, with optional vision and audio encoders listed at 170 million and 300 million parameters. Google also says developers can shorten output vectors from 768 dimensions to 512, 256 or 128 using Matryoshka Representation Learning, reducing storage and memory use. The company estimates up to sixfold storage reduction, depending on the chosen vector size and implementation.

At a glance
announcementWhen: Announced Oct. 6, 2026
The developmentGoogle released EmbeddingGemma 2, a multimodal embedding model intended for local search and retrieval on consumer hardware.

Local Search Across Media Types

The release targets developers building search, routing and retrieval features for devices that may have limited or intermittent connectivity. If embeddings are generated locally, applications can avoid uploading source material for that processing step, and may reduce the delay associated with sending data to a server. Google presents these as benefits of on-device operation; actual privacy protections depend on how an application stores, uses and transmits data elsewhere in its pipeline.

A shared representation for several media types could also make it easier to retrieve related material across formats—for example, matching spoken notes with video or images. That could be useful in personal media libraries, field-work tools and local document systems. The model does not by itself provide a complete search product: developers still need an indexing and retrieval system, and must test whether results meet their accuracy and speed requirements.

For applications that generate answers from retrieved material, Google says EmbeddingGemma 2 can be paired with Gemma 4 in an on-device retrieval-augmented generation pipeline. Google notes that the models share a text tokenizer and audio encoder, which it says can reduce combined memory requirements. The announcement does not give a single total memory figure for such a combined system.

Amazon

on-device multimodal embedding model

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Text Embeddings to Multimodal Retrieval

Embedding models turn content into numerical vectors so software can find items with similar meaning, rather than relying only on exact keyword matches. Google introduced the original EmbeddingGemma as a lightweight text model for organizing and searching information on consumer hardware. The new release broadens that goal to multiple content types and code.

Google reports that EmbeddingGemma 2 improves its predecessor’s result on the MTEB Code benchmark from 68.76 to 78.68, a 9.92-point increase. It also says the model scores well among multimodal embedding models with fewer than one billion parameters across MTEB Code and the Massive Audio Embedding Benchmark. The source announcement points readers to a model card for full evaluation metrics; it does not reproduce all benchmark settings or provide independent evaluations.

For deployment, Google lists model weights on Hugging Face and Kaggle, with availability in Gemini Enterprise Agent Platform Model Garden described as coming soon. It also names Google AI Edge MediaPipe and LiteRT, browser options including transformers.js and WebGPU, and serving tools such as Transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio. These are stated integration options, not evidence that every configuration has the same performance.

““EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings.””

— Google, in its Oct. 6, 2026 announcement

Amazon

multimedia search and retrieval device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests and Device Costs

The announcement does not include independent benchmark results, detailed evaluation settings, or comparisons that can be checked against the full model card. Google’s claims about leading performance and quality per parameter should be read as vendor claims until other researchers or developers report comparable tests.

Google gives a device-specific memory estimate: with quantization, the model uses about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal model on a Pixel 11 Pro. The source does not specify the full range of hardware, latency, energy use, or performance across tasks. It is also not clear how consistently the listed context capacities translate into accurate retrieval, or what trade-offs result from shorter embeddings.

Amazon

local audio and video search tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Model Access and Developer Testing

Developers can access the weights through Hugging Face and Kaggle and test deployment with the tools listed by Google. The company says Model Garden availability is planned for a later date but does not give a specific release date. Google also directs developers to its documentation, inference and fine-tuning guides, and a LiteRT guide for building on-device search and retrieval systems.

The next useful evidence will come from model-card details and tests on the range of devices developers intend to support. Those tests can establish how model quality, memory use, speed and battery consumption vary by modality and application. Until those results are available, the release confirms the model’s stated design and availability channels, but not uniform performance in real-world products.

Amazon

multimodal content search hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is EmbeddingGemma 2?

It is Google’s 740-million-parameter embedding model for representing text, code, images, audio and video in a shared space used by search and retrieval systems.

Can it run without an internet connection?

Google says the model is designed for on-device inference, which can support offline workflows. Whether a finished app works fully offline depends on its other components and how developers configure it.

What license does Google use?

Google says EmbeddingGemma 2 is released under the Apache 2.0 license. Developers should consult the model’s license and documentation for applicable terms.

Where can developers get the model?

Google lists Hugging Face and Kaggle as sources for the weights. The company says Gemini Enterprise Agent Platform Model Garden availability is coming soon, without specifying a date.

Are the performance claims independently verified?

The announcement reports Google’s benchmark results and refers to a model card, but provides no independent evaluation. Third-party testing will be needed to compare performance across hardware and use cases.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will OpenAI Release GPT-5.6 Before Jul 7, 2026?

Market activity suggests OpenAI may release GPT-5.6 before July 2026, but official confirmation is pending. Key details and uncertainties explained.

Cloud’s Hidden Memory Bill

Rising DRAM prices are increasing cloud costs subtly, leading to higher bills and potential shifts in infrastructure strategies amid ongoing shortages.

Exploring Anthropic Model Availability On Amazon Bedrock In Seoul And Singapore

Anthropic says its models on Amazon Bedrock can run inference in Seoul and Singapore; model coverage, timing and data-handling details are not specified.

The Perception Trap: When AI All Reads From The Same Script

Exploring how widespread reliance on a few AI models is creating a collective perception bias, risking market instability and societal brittleness.