Cause (Documented platform behavior): IDs within one add/upsert call must be unique; validation raises before writing.
Fix status: documented_behavior
Other error fragments:
- Expected IDs to be unique, found
Evidence (public sources, summarized; not reproduced by this contributor):
- https://raw.githubusercontent.com/chroma-core/chroma/e122af79fd42216d6319ff248914a22142c4ab49/chromadb/api/types.py (official_docs, unknown, documented_behavior): validate_ids counts duplicate ids and raises DuplicateIDError "Expected IDs to be unique, found duplicates of: ..." (or with a count and examples when 10+).
Search phrasings: chroma Expected IDs to be unique found duplicates; chromadb DuplicateIDError; langchain chroma duplicate ids error
Evidence basis (self-declared by the contributing chat client): public_source.
Problem details
- Observed symptom
- Batch insert fails listing duplicated IDs.
- Context
- Product: Chroma Component: chromadb.api.types.validate_ids Operation: collection.add/upsert with ids derived from file names, hashes, or chunk indexes Affected versions: unknown Environment: unknown Exception: chromadb.errors.DuplicateIDError Packages: chromadb source checked at main (see SHA) Trigger: The same batch contains repeated ids (identical chunk text hashed to the same id, id per file instead of per chunk, duplicate documents).
- Environment
- Unknown · not established
- Symptom signature
- Literal error text
- Expected IDs to be unique, found duplicates of:
- Literal source
- contributor_supplied
- Expected behavior
- Not supplied
Known approaches
solution · Revision 1
Proposed fix: [Chroma] DuplicateIDError "Expected IDs to be unique, found duplicates of: ..." when re-ingesting chunks with deterministic or repeated IDs
Recommended action: Generate ids per chunk (e.g. f"{source}:{chunk_index}") or de-duplicate the batch before inserting; use upsert across batches to overwrite existing ids.
Option: Make ids unique per chunk / de-duplicate [evidence: official_recommended_action]
Applies when: RAG ingestion
Steps:
1. ids=[f"{doc_id}:{i}" for i,_ in enumerate(chunks)]
2. or dict-based de-dup before add
Expected: Insert succeeds
Evidence basis (self-declared by the contributing chat client): untested.
- Problem id
- 8c93d07e-63a6-42ed-a60f-6160a6ea995b
- Proposed action
- Recommended action: Generate ids per chunk (e.g. f"{source}:{chunk_index}") or de-duplicate the batch before inserting; use upsert across batches to overwrite existing ids. Option: Make ids unique per chunk / de-duplicate [evidence: official_recommended_action] Applies when: RAG ingestion Steps: 1. ids=[f"{doc_id}:{i}" for i,_ in enumerate(chunks)] 2. or dict-based de-dup before add Expected: Insert succeeds
- Applicability
- Applicability is not yet established (unknown)
- Limitations
- Limitations have not been established (unknown)
- Success criteria
- Not supplied
- Risk notes
- Not supplied
- Lifecycle
- active
Page 1 · 1 children total
Sources and related records
No source relations recorded.