Should Clipboard Apps Cluster Items With Embeddings?
Embedding-based clustering is one of the more attractive AI features for a clipboard manager. A 200-item history is hard to scan; ten thematic groups (URLs, code snippets, error logs, chat messages, file paths, addresses, phone numbers, dates, financial data, other) are easy. The question is not whether clustering is useful — it is — but where the embeddings run. Local embeddings keep the clipboard private; vendor API embeddings send every copy to a third party. This guide covers the design, the trade-offs, and what a local-first version would require. For related reading, see Windows Recall, Click to Do, and copied text, file-transfer shelves vs official Nearby Share, what local-first software means for clipboard apps, and why some clipboard apps will never sync.
What embedding clustering actually does
The pipeline is:
- Embed each item. Take the clipboard item's text, run it through an embedding model, get a vector (typically 384 to 1536 dimensions). The vector represents the item's meaning in a way that similar items have similar vectors.
- Cluster the vectors. Run a clustering algorithm (k-means, HDBSCAN, agglomerative) over the vectors. The output is a set of groups, where items in the same group are semantically similar.
- Label the groups. Either let the user name them, or use a small model to generate labels ("URLs," "code snippets," "error logs").
- Display. Show the groups in the clipboard manager's UI, instead of (or in addition to) the chronological list.
Each step has a cost. Embedding is the most expensive; a 384-dimension local model takes 50–200 ms per item on a typical laptop, which is fine for a single item but adds up for a 500-item history re-clustered on every launch. Clustering is cheap; k-means on 500 vectors is sub-second. Labelling is expensive if a model is used, cheap if the user does it. Display is the easiest part.
The two embedding paths
Local embeddings
Local embeddings run on the user's device. The clipboard content never leaves the machine. The user installs an embedding model (sentence-transformers, Ollama embeddings, Foundry Local embeddings, or a bundled model), and the clustering uses that model.
The privacy posture is the same as the clipboard manager itself: data stays on the device, the user controls it. The cost is RAM (an embedding model needs 200–500 MB resident) and the setup friction of installing a model.
Reasonable local embedding models in 2026:
all-MiniLM-L6-v2(sentence-transformers) — 384 dimensions, ~80 MB, fast on CPU. Good default.bge-small-en-v1.5— 384 dimensions, ~130 MB, slightly better quality.nomic-embed-text(via Ollama) — 768 dimensions, ~270 MB, good quality.- Foundry Local embeddings — Microsoft's local embedding runtime, bundled model.
For a clipboard manager, all-MiniLM-L6-v2 is the obvious starting point: small, fast, free, and good enough for the kind of clustering a clipboard needs.
Vendor API embeddings
Vendor API embeddings send the clipboard item's text to a vendor's API, get a vector back, and use that vector for clustering. Examples:
- OpenAI embeddings (
text-embedding-3-small,text-embedding-3-large). - Cohere embeddings.
- Google Vertex AI embeddings.
- Azure OpenAI embeddings.
The privacy posture is the same as for any remote model: the vendor sees the text, retains it per its policy, and may use it for training unless the user opts out. For a clipboard, this means every copied item — passwords, API keys, private messages, financial data — is sent to the vendor.
This is the leak. A clipboard manager that calls a vendor embedding API on every copy is, in the worst case, sending everything the user copies to a third party. The clustering UI may be pretty; the privacy cost is not.
How to tell which is which
The same questions as for AI clipboard managers in general:
- Where does the embedding model run? "On-device" or "locally" means local. "Powered by OpenAI" or "uses Cohere" means remote.
- Does it work offline? Disconnect from the internet. If clustering still works, the model is local. If it stops, the model is remote.
- What does the network monitor show? See telemetry-free desktop utilities: how to check. Large POST requests to an embedding API on every copy are the signal.
- What does the privacy policy say? Vendors that use your data for training will say so.
For apps with source available, grep for embedding endpoints:
grep -r "openai.com/v1/embeddings" src/
grep -r "cohere.ai" src/
grep -r "generativelanguage.googleapis.com" src/
The honest case for each posture
Local-only
Best for: every clipboard user. The privacy cost of remote embeddings is too high for the benefit of slightly better clustering.
Cost: 80–500 MB of RAM for the embedding model, plus the engineering work to integrate it.
Honest summary: this is the only acceptable default for a clipboard manager. Anything else sends secrets to a vendor.
Remote, opt-in per item
Best for: users who want to cluster a specific subset of items (a research session, a project's worth of snippets) and are willing to send those items to a vendor.
Cost: the user has to remember to opt in, and the vendor's data retention still applies.
Honest summary: this is a reasonable compromise for a power-user feature. It should not be the default.
Remote, always-on
Best for: nobody. This is the same leak as remote-always-on AI summarisation, applied to embeddings.
Honest summary: do not install a clipboard manager that does this.
What a real local-first implementation would require
A clipboard manager shipping local embedding clustering would need to:
- Bundle or download an embedding model.
all-MiniLM-L6-v2is small enough to bundle. Larger models should be opt-in downloads. - Embed on copy, not on display. Computing the embedding when the item is copied means the cost is amortised; clustering at display time is then fast.
- Cache embeddings. Store the vector alongside the item. Re-embedding a 500-item history on every launch is wasteful.
- Re-cluster incrementally. When a new item is added, assign it to an existing cluster or create a new one. Full re-clustering is only needed when the user asks for it.
- Let the user label clusters. Auto-labels from a small local model are nice but not necessary. User labels are more accurate.
- Make it optional. Some users do not want clustering. The default should be the chronological list, with clustering as an opt-in view.
This is a real feature, not a weekend project. The engineering cost is meaningful: model integration, vector storage, clustering algorithm, UI work. For a solo-maintained project, this is a multi-month feature, not a quick win.
Where Edge-Drop sits
Edge-Drop does not ship embedding clustering. There is no shipped clustering, no remote embedding API integration, and no local embedding model integration. The honest statement is "no clustering," not "clustering coming soon." Users who want clustering on top of Edge-Drop can run a local embedding model alongside it and build a script that reads from Edge-Drop's history, but this is a user-built integration, not a product feature.
This is not a permanent stance. It is a scoping decision: clustering introduces a privacy threat model and an engineering cost that the rest of the product does not have, and shipping it carelessly would make the product worse. For more on the scope decision, see what Edge-Drop is not trying to become.
For the broader AI-on-the-clipboard question, see AI clipboard managers: useful or a new leak and PowerToys Advanced Paste AI: local vs cloud models.
What to look for in 2026
The category is moving. Signals worth tracking:
- OS-level embedding APIs. Windows Copilot Runtime may eventually ship a local embedding API that clipboard managers can call without bundling a model. As of mid-2026, this is not the default; check the current SDK.
- Smaller, faster embedding models. The state of the art moves quickly; a 50 MB model with 384-dimension embeddings and good quality is plausible within a year.
- Vendor data retention policies. Vendors may ship "we do not retain embedding API inputs" policies. Verify the policy, do not trust the marketing.
- On-device sensitive-data filtering. A small local model that flags API keys and passwords before they are embedded is a real defence. The combination "filter, then embed only non-sensitive items" is the right posture for a remote embedding path, if one is used at all.
For most users, the practical takeaway is: prefer tools that do not call remote embedding APIs on every copy. If a tool ships clustering, verify the model is local. If it is not, treat the tool as a leak. For more on the broader local-first argument, see what local-first software means for clipboard apps.
Related reading
- File-Transfer Shelves vs Official Nearby Share
- Why Tray Icons Lost to Edges and Launchers
- What Local-First Software Means for Clipboard Apps
- How to Enable Clipboard History in Windows 11
Sources
- Hugging Face — Sentence Transformers — community-maintained library of embedding models, including
all-MiniLM-L6-v2, the default for local-first embedding features - Ollama — Embeddings documentation — official documentation for using Ollama as a local embedding backend
- Microsoft Learn — Foundry Local embeddings — official documentation for Microsoft's local embedding runtime
- OpenAI — Embeddings API documentation — official documentation for the OpenAI embeddings API, useful for understanding what a remote embedding path sends
- OWASP — AI security and privacy guide — community reference for the privacy considerations specific to AI features in end-user software
Deepender Yadav is a B.Tech Computer Science Engineering student and software developer interested in building practical software and open-source projects.
GitHub · LinkedInCopy. Stack. Drop.
Transform your clipboard into an interactive edge shelf. Stack, pin, and drag assets into any app with zero friction.
Download for Windows Get from Microsoft Store
How to Install Guide · First 10 Minutes Guide · Drag & Drop Guide · Edge-Drop vs Win+V · Support
Free · Lightweight · Privacy First