Here is a fact most business owners never hear: before an AI can find the right answer in your files, it has to turn every sentence, photo and voice note into a long list of numbers. Not a summary. Not a set of keywords. A list of several hundred numbers that captures what the content means.
That list is called an embedding, and it quietly powers almost every "smart search" and "ask your documents" feature you have seen in the last two years. On 6 October 2026, Google DeepMind released EmbeddingGemma 2, a free model that can create embeddings for text, images, audio and video on an ordinary phone. The term is about to show up in a lot more sales pitches.
This guide explains what an embedding is, how it works, where it already helps small businesses and what to ask before you pay for anything that uses one.
What is an embedding in AI?
An embedding is a list of numbers that represents the meaning of a piece of content, such as a sentence, a product photo or a short recording. Content with similar meaning gets similar numbers. This lets a computer judge that "Do you deliver to Tsuen Wan?" and "Can you ship to the New Territories?" are close in meaning, even though they share almost no words.
Think of it as giving every piece of content a precise address on a giant map of meaning. Questions about delivery live in one neighbourhood. Questions about refunds live in another. Photos of red dresses sit near other photos of red dresses, and not far from the words "red dress".
The AI does not "read" your documents the way a person does each time you search. It compares addresses on this map, which is why it can be fast even with thousands of files.
How does an embedding work, step by step?
An embedding works in three steps: a model converts each piece of your content into a list of numbers once, those lists are stored in a searchable index, and every new question is converted the same way and matched to the closest stored items. The closest matches are returned as results or handed to a chatbot to write an answer.
In practice, the flow looks like this:
--- Step 1, convert: An embedding model reads each item, for example every page of your price list or every past customer email, and outputs a list of numbers for it. EmbeddingGemma 2 produces 768 numbers per item by default.
--- Step 2, store: The lists go into a database built for this job, often called a vector database. "Vector" is simply the technical word for a list of numbers.
--- Step 3, compare: When someone asks a question, it is converted into numbers too. The system finds the stored items whose numbers are closest, which means closest in meaning.
A useful comparison is a wine shop with a very experienced manager. You describe what you liked last time, and the manager walks straight to the right shelf without reading every label. The embedding is the manager's sense of where each bottle belongs.
What changed with EmbeddingGemma 2 in October 2026?
EmbeddingGemma 2, released by Google DeepMind on 6 October 2026, places text, code, images, video and audio on the same map of meaning. It is free to download under the Apache 2.0 licence, and Google says the full version runs in about 567 MB of memory on a recent phone, with a text-only version using about 191 MB.
According to Google's developer announcement, a single model can now match a voice memo to a video clip, or a typed question to a photo. Earlier versions handled text only. The original EmbeddingGemma, released in 2025, passed 20 million downloads, according to Google.
Three details matter for business owners:
--- It runs locally. Content can be indexed and searched on the device, so files do not have to be sent to an outside service just to be searched.
--- It is modular. Developers can load only the text part, or add the image and audio parts when needed, which keeps small setups lightweight.
--- It has no built-in safety filter. The model card states it is a pre-trained model without safety tuning, so whoever builds the product is responsible for filtering and testing.
You do not need to download anything yourself. The practical effect is that vendors can now offer "search by meaning" across photos and recordings at a lower cost, and some will offer it without moving your data off your own machines.
How can a Hong Kong SME use embeddings today?
A Hong Kong SME uses embeddings whenever it lets customers or staff search by meaning instead of exact words. Common examples are a website search that understands casual questions, a staff assistant that finds the right clause in a policy manual, and a product catalogue where shoppers can upload a photo to find similar items.
Some concrete scenarios:
--- Online shop: A customer types "something warm for a Hokkaido trip" and gets down jackets and thermal wear, even though no product page uses the word "Hokkaido".
--- Clinic or salon: A front-desk assistant finds the right aftercare instructions when a customer describes symptoms in their own words, in English or Chinese.
--- Trading company: Staff search ten years of supplier emails for "the factory that had the packaging problem" without remembering the company name.
--- Property agency: Agents find past listings that match a buyer's description, such as "quiet, near MTR, good for a family with a dog".
--- Customer service: Incoming WhatsApp questions are matched to the closest approved answer, so replies stay consistent across staff.
None of these require a new app on the customer's side. The embedding work happens behind the search box or chat window your customers already use, which is why most owners have used embeddings without knowing the word.
The quality of the result depends on the quality of what you feed in. If your price list is out of date, embeddings will find the out-of-date price very efficiently.
What do people get wrong about embeddings?
The most common mistake is assuming embeddings make AI accurate. Embeddings make AI good at finding related content, but related is not the same as correct. A system can retrieve a passage that is similar in meaning yet outdated, or miss an exact figure such as an invoice number that keyword search would have caught instantly.
Other misconceptions worth clearing up:
--- "Embeddings replace keyword search." They work best together. Product codes, phone numbers and names are often found more reliably by exact matching. Many good systems use both, an approach called hybrid search.
--- "Embeddings are a summary of my document." They are not readable. You cannot open an embedding and see your text, although researchers have shown that some content can be partly reconstructed, so treat embeddings of sensitive files as sensitive data.
--- "One embedding model works for everything." Models differ in language strength. Test with real Cantonese and mixed English-Chinese questions from your own customers before you commit.
--- "Once set up, it stays accurate." If you change your prices or policies, the stored embeddings must be refreshed. Stale indexes are a quiet source of wrong answers.
How do you check an AI search tool before paying for it?
You check an AI search tool by testing it with your own messy, real questions and measuring how often the right document appears in the top three results. Ask the vendor where embeddings are stored, how often they are refreshed after content changes, and whether keyword matching is used alongside meaning-based search for codes and names.
A simple five-step test any owner can run:
--- Collect 20 real questions from customer messages or staff, including typos, slang and mixed languages.
--- Write down the correct answer source for each, such as the right page or policy section.
--- Run all 20 through the tool and count how many times the correct source appears in the top three results.
--- Change one price or policy and check how long the tool takes to reflect it.
--- Ask about data location. Find out whether your content and its embeddings stay on your device, in Hong Kong, or overseas, which matters under the Personal Data (Privacy) Ordinance. Our explainer on data residency covers the questions to ask.
If the tool gets 17 or more right out of 20, it is likely doing the search step well. Fewer than 12 is a sign to keep looking.
Embeddings FAQ
Quick answers to the questions business owners ask most about embeddings, vector databases and meaning-based search.
Is an embedding the same as a chatbot like ChatGPT?
No. An embedding model finds and ranks content by meaning, while a chatbot model writes new sentences. Many business AI tools use both: the embedding model first finds the three or four most relevant passages in your files, and the chatbot then writes a reply based on those passages. This combination is usually called retrieval-augmented generation, or RAG.
The split matters when you evaluate a product. If an AI assistant gives wrong answers about your own policies, the problem is often in the search step, not the writing step. We explain the writing-side risk in our guide to why AI confidently makes things up.
Do I need a vector database to use embeddings?
Not directly. Most business tools that offer AI search include the storage behind the scenes. You only choose a vector database yourself if you are building a custom system with a developer.
Are embeddings expensive?
Usually not. Creating embeddings is one of the cheapest AI tasks, and free models such as EmbeddingGemma 2 can run on existing hardware. The larger costs tend to be setup, data cleaning and the chatbot that writes answers.
Do embeddings work with Chinese and Cantonese?
Many modern models support Chinese, but quality varies, especially for written Cantonese and mixed-language messages. Always test with real customer messages before relying on one.
Can embeddings search photos and voice messages?
Yes, with a multimodal model. EmbeddingGemma 2 places images, audio and video on the same map as text, so a typed question can find a matching photo or recording.
Conclusion: the map behind smart search
An embedding is the AI's map of meaning. It is the reason a search box can understand "something warm for a Hokkaido trip", and the reason a staff assistant can find the right paragraph in a 200-page manual. With free, phone-sized models like EmbeddingGemma 2, this capability is becoming cheaper and easier to run close to your own data.
The owner's job is unchanged: keep the source content accurate, test with real questions and ask where the data lives. Get those right, and meaning-based search becomes one of the most useful AI upgrades a small business can make. We understand AI. UD stands with you.
Reviewed by the UD AI team.
Meaning-based search is only as good as the content and setup behind it. Talk to us about AI staff that search your own price lists, manuals and past messages, and we will walk you through it step by step, from preparing your documents to testing with real customer questions.