Teradata Package for LangChain Function Reference - add_datasets - Teradata® Package for LangChain - Look here for syntax, methods and examples for the functions included in the Teradata langchain-teradata package.
Teradata® Package for LangChain Function Reference
- Deployment
- VantageCloud
- Edition
- Enterprise
- Product
- Teradata® Package for LangChain
- Release Number
- 20.00.00.01
- Published
- December 2025
- ft:locale
- en-US
- ft:lastEdition
- 2025-12-19
- dita:id
- Langchain-Teradata_FxRef_Lake
- Product Category
- Teradata Vantage
- libs.teradata.langchain_teradata.TeradataVectorStore.add_datasets = add_datasets(self, data, **kwargs)
- DESCRIPTION:
Adds the specified data to an existing content-based vector store.
Creates a new Vector Store in case it does not exists.
PARAMETERS:
data:
Required Argument.
Specifies the table name(s)/teradataml DataFrame(s) to be indexed or
added to the vector store. Teradata recommends to use teradataml DataFrame as input.
Notes:
* Multiple tables/views/DataFrames can be passed in "dataset".
* If the table is in another database than
the database in use, make sure to pass in
the fully qualified name or a DataFrame object.
For example,
* If the table name is 'amazon_reviews' and it is
under 'oaf' database which is not the user's
logged in database
* Pass argument as DataFrame(in_schema('oaf', 'amazon_reviews'))
* If multiple tables/views are passed, each table should
have the columns which are mentioned in "data_columns"
and "key_columns".
* When "target_database" is not set and only table name is passed to
"dataset", then the input is searched in default database.
Types: str or list of str or DataFrame
name:
Optional Argument.
Specifies the name of the vector store to be created from
the input dataset(s).
Types: str
data_columns:
Required Argument.
Specifies the name(s) of the data column(s) to be used
for embedding generation(vectorization).
Note:
When multiple data columns are specified, data is unpivoted
to get a new key column "AttributeName" and a single data column
"AttributeValue".
Types: str or list of str
key_columns:
Optional Argument.
Specifies the name(s) of the key column(s) to be used for indexing.
Types: str, list of str
description:
Optional Argument.
Specifies the description of the vector store.
Types: str
target_database:
Optional Argument.
Specifies the database name where the vector store is created.
Notes:
* If not specified, vector store is created in the database
which is in use.
Types: str
vector_column:
Optional Argument.
Specifies the name of the column to be used for storing
the embeddings.
Default Value: vector_index
Types: str
metadata_columns:
Optional Argument.
Specifies the list of input column names to be used for metadata.
These columns just get accumulated in the vector store.
Types: list[str]
metadata_descriptions:
Optional Argument.
Specifies the deescriptions of the metadata columns. One value for each metadata column.
Note:
Applicable to all store types except metadata-based store type.
Types: list[str]
use_simd:
Optional Argument.
Specifies whether to use SIMD for faster processing.
Types: bool
embedding_datatype:
Optional Argument.
Specifies the data type of the embeddings to be used.
Default Value: VECTOR32
Permitted Values: VECTOR32, VECTOR64
Types: str
embeddings_dims:
Required for NVIDIA NIM, Optional otherwise.
Specifies the number of dimensions to be used for generating the embeddings.
The value depends on the "embeddings".
Note:
* Default dimesions is set to 1024 for embedding-based vector store.
Default Value:
For AWS:
* amazon.titan-embed-text-v1: 1536
* amazon.titan-embed-image-v1: 1024
* amazon.titan-embed-text-v2:0: 1024
For Azure:
* text-embedding-ada-002: 1536
* text-embedding-3-small: 1536
* text-embedding-3-large: 3072
Permitted Values:
*For AWS:
* amazon.titan-embed-text-v1: 1536
* amazon.titan-embed-image-v1: [256, 384, 1024]
* amazon.titan-embed-text-v2:0: [256, 512, 1024]
*For Azure:
* text-embedding-ada-002: 1536 only
* text-embedding-3-small: 1 <= dims <= 1536
* text-embedding-3-large: 1 <= dims <= 3072
Types: str
chat_completion_max_tokens:
Required for NVIDIA NIM, Optional otherwise.
Specifies the maximum number of tokens to be generated by the
"chat_completion_model".
Default Value: 16384
Permitted Values: [1, 16384]
Types: int
model_urls:
Optional Argument.
Specifies the URL and models to be used for embedding, chat completion
and guardrails.
Note:
* Refer to the ModelUrlParams class for more details.
Types: ModelUrlParams
update_style:
Optional Argument.
Specifies the style to be used for alter operation of the data
from the vector store when "search_algorithm" is KMEANS/HNSW.
Default Value: MINOR
Permitted Values: MINOR, MAJOR
Types: str
metadata_operation:
Optional Argument.
Specifies the operation to be performed on metadata columns
during update.
- ADD - add new metadata columns
- DELETE - remove existing metadata columns
- MODIFY - change the description of the existing metadata columns
Note:
* Applicable to all store types except the METADATA-BASED store type.
Default Value: ADD
Permitted Values: ADD, DELETE, MODIFY
Types: str
embedding:
Required for NVIDIA NIM, Optional otherwise.
Specifies the embeddings model to be used for generating the
embeddings.
Default Value:
For AWS: amazon.titan-embed-text-v2:0
For Azure: text-embedding-3-small
Permitted Values:
For AWS:
* amazon.titan-embed-text-v1
* amazon.titan-embed-image-v1
* amazon.titan-embed-text-v2:0
For Azure:
* text-embedding-ada-002
* text-embedding-3-small
* text-embedding-3-large
Types: str, TeradataAI, LangChain Embeddings
chat_completion_model:
Required for NVIDIA NIM, Optional otherwise.
Specifies the name of the chat completion model to be used for
generating text responses.
Default Value:
For AWS: anthropic.claude-3-haiku-20240307-v1:0
For Azure: gpt-35-turbo-16k
Permitted Values:
*For AWS:
* anthropic.claude-3-haiku-20240307-v1:0
* anthropic.claude-instant-v1
* anthropic.claude-3-5-sonnet-20240620-v1:0
*For Azure:
* gpt-35-turbo-16k
Types: str, TeradataAI, LangChain BaseChatModel
RETURNS:
None
RAISES:
TeradataMlException
EXAMPLES:
>>> from langchain_teradata import TeradataVectorStore
>>> from teradataml import DataFrame
>>> from teradatagenai import load_data
>>> load_data("byom", "amazon_reviews_25")
>>> amazon_reviews_25 = DataFrame('amazon_reviews_25')
>>> amazon_reviews_10 = load_data("byom", "amazon_reviews_10")
# Example 1: Create an instance of an 'content-based' vector store by passing the 'amazon_reviews_25' table.
>>> vs_instance1 = TeradataVectorStore.from_datasets(name = "vs_example_1",
data = "amazon_reviews_25",
data_columns = ["rev_text"],
embedding = "amazon.titan-embed-text-v1")
# Add data to an existing content-based vector store "vs_example_1"
>>> vs_instance1.add_datasets(data=amazon_reviews_10)
# Example 2: Create a new 'content-based' vector store by passing the 'amazon_reviews_10' table.
>>> vs_instance2 = TeradataVectorStore()
>>> vs_instance2.add_datasets(name = "vs_example_2",
data = "amazon_reviews_10",
data_columns = ["rev_text"],
embedding = "amazon.titan-embed-text-v1")