Teradata Package for LangChain Function Reference - from_documents - Teradata® Package for LangChain - Look here for syntax, methods and examples for the functions included in the Teradata langchain-teradata package.

Teradata® Package for LangChain Function Reference

Deployment
VantageCloud
Edition
Enterprise
Product
Teradata® Package for LangChain
Release Number
20.00.00.01
Published
December 2025
ft:locale
en-US
ft:lastEdition
2025-12-19
dita:id
Langchain-Teradata_FxRef_Lake
Product Category
Teradata Vantage
libs.teradata.langchain_teradata.TeradataVectorStore.from_documents = from_documents(name, documents, embedding=None, **kwargs) class method of libs.teradata.langchain_teradata.vector_store.TeradataVectorStore
DESCRIPTION:
    Creates a new vector store, either 'file-based' or 'content-based',
    depending on the type of input documents.
    If the input is PDF file(s) or file path(s), a file-based vector store is created.
    If the input is LangChain Document object(s), a content-based vector store
    is created with metadata stored in "metadata_columns"
    Raises an error if a vector store with the specified name already exists.
    Notes:
        * If Document objects contain metadata keys that are Teradata reserved keywords,
          users can rename those keys using the `rename_metadata_keys` function.
        * Only admin users can use this method.
        * Refer to the 'Admin Flow' section in the
          User Guide for details.
 
PARAMETERS:
    name:
        Required Argument.
        Specifies the name of the vector store.
        Type: str
 
    documents:
        Required Argument.
        Specifies the input dataset of document files or LangChain Document objects.
        For files:
            - Accepts a directory path, wildcard pattern or a list of file paths
            - Only PDF format is supported.
            - Files are processed and stored as chunks in a database table.
        For LangChain Document objects:
            - Accepts a list of Document objects
            - Documents are processed and stored as chunks in a database table.
        Notes:
            * Input can be either file(s)/file path(s) or LangChain Document
              objects, not both.
        Examples:
            Example 1: Multiple files specified within a list
            >>> documents = ['file1.pdf', 'file2.pdf']
 
            Example 2: Path to the directory containing PDF files 
            >>> documents = "/path/to/pdfs"
 
            Example 3: Path to directory containing PDF files as a wildcard string
            >>> documents = "/path/to/pdfs/*.pdf"
 
            Example 4: Path to directory containing PDF files and subdirectories of PDF files
            >>> documents = "/path/to/pdfs/**/*.pdf"
 
            Example 5: List of LangChain Document objects
            >>> from langchain_core.documents import Document
            >>> documents = [Document(page_content="This is a test document", id="doc1"),
                             Document(page_content="This is another test document", id="doc2")]
        Types: str, list, LangChain Document object
 
    object_names:
        Optional Argument.
        Specifies the table name to store file content splits.
        Notes:
            * Applicable only for file-based inputs.
            * Only one table name should be specified.
        Type: str
 
    target_database:
        Optional Argument.
        Specifies the database name where the vector store and file content
        splits are created/stored.
        Note:
            If not specified, uses the current database.
        Type: str
 
    data_columns:
        Optional Argument.
        Specifies the column name(s) to store the content splits.
        Notes:
            * Applicable only for file-based inputs.
        Type: str
 
    embedding:
        Required for NVIDIA NIM, Optional otherwise.
        Specifies the embeddings model to be used for generating the
        embeddings.
        Default Value:
            For AWS: amazon.titan-embed-text-v2:0
            For Azure: text-embedding-3-small
        Permitted Values:
            For AWS:
                * amazon.titan-embed-text-v1
                * amazon.titan-embed-image-v1
                * amazon.titan-embed-text-v2:0
            For Azure:
                * text-embedding-ada-002
                * text-embedding-3-small
                * text-embedding-3-large
        Types: str, TeradataAI, LangChain Embeddings
 
    chat_completion_model:
        Required for NVIDIA NIM, Optional otherwise.
        Specifies the name of the chat completion model to be used for
        generating text responses.
        Default Value:
            For AWS: anthropic.claude-3-haiku-20240307-v1:0
            For Azure: gpt-35-turbo-16k
        Permitted Values:
            *For AWS:
                * anthropic.claude-3-haiku-20240307-v1:0
                * anthropic.claude-instant-v1
                * anthropic.claude-3-5-sonnet-20240620-v1:0
            *For Azure:
                * gpt-35-turbo-16k
        Types: str, TeradataAI, LangChain BaseChatModel
 
    description:
        Optional Argument.
        Specifies the description of the vector store.
        Types: str
 
    vector_column:
        Optional Argument.
        Specifies the name of the column to be used for storing
        the embeddings.
        Default Value: vector_index
        Types: str
 
    metadata_columns:
        Optional Argument.
        Specifies the list of input column names to be used for metadata.
        These columns just get accumulated in the vector store.
        Types: list[str]
 
    metadata_descriptions:
        Optional Argument.
        Specifies the deescriptions of the metadata columns. One value for each metadata column.
        Note:
            Applicable to all store types except metadata-based store type.
        Types: list[str]
 
    use_simd:
        Optional Argument.
        Specifies whether to use SIMD for faster processing.
        Types: bool
 
    embedding_datatype:
        Optional Argument.
        Specifies the data type of the embeddings to be used.
        Default Value: VECTOR32
        Permitted Values: VECTOR32, VECTOR64
        Types: str
 
    embeddings_dims:
        Required for NVIDIA NIM, Optional otherwise.
        Specifies the number of dimensions to be used for generating the embeddings.
        The value depends on the "embeddings".
        Note:
            * Default dimesions is set to 1024 for embedding-based vector store.
        Default Value:
            For AWS:
                * amazon.titan-embed-text-v1: 1536
                * amazon.titan-embed-image-v1: 1024
                * amazon.titan-embed-text-v2:0: 1024
            For Azure:
                * text-embedding-ada-002: 1536
                * text-embedding-3-small: 1536
                * text-embedding-3-large: 3072
        Permitted Values:
            *For AWS:
                * amazon.titan-embed-text-v1: 1536
                * amazon.titan-embed-image-v1: [256, 384, 1024]
                * amazon.titan-embed-text-v2:0: [256, 512, 1024]
            *For Azure:
                * text-embedding-ada-002: 1536 only
                * text-embedding-3-small: 1 <= dims <= 1536
                * text-embedding-3-large: 1 <= dims <= 3072
        Types: str
 
    chat_completion_max_tokens:
        Required for NVIDIA NIM, Optional otherwise.
        Specifies the maximum number of tokens to be generated by the
        "chat_completion_model".
        Default Value: 16384
        Permitted Values: [1, 16384]
        Types: int
 
    model_urls:
        Optional Argument.
        Specifies the URL and models to be used for embedding, chat completion
        and guardrails.
        Note:
            * Refer to the ModelUrlParams class for more details.
        Types: ModelUrlParams
 
    chunk_size:
        Optional Argument.
        Specifies the number of characters in each chunk to be used while
        splitting the input file.
        Note:
            Applicable only for 'file-based' vector stores.
        Default Value: 512
        Types: int
 
    optimized_chunking:
        Optional Argument.
        Specifies whether an optimized splitting mechanism supplied by
        Teradata should be used. The documents are parsed internally in an
        intelligent fashion based on file structure and chunks are dynamically
        created based on section layout.
        Notes:
            * The "chunk_size" field is not applicable when
              "optimized_chunking" is set to True.
            * Applicable only for 'file-based' vector stores.
        Types: bool
 
    header_height:
        Optional Argument.
        Specifies the height (in points) of the header section of a PDF
        document to be trimmed before processing the main content.
        This is useful for removing unwanted header information
        from each page of the PDF. Recommended value is 55.
        Note:
            * Applicable only for 'file-based' vector stores.
        Types: int
 
    footer_height:
        Optional Argument.
        Specifies the height (in points) of the footer section of a PDF
        document to be trimmed before processing the main content.
        This is useful for removing unwanted footer information from
        each page of the PDF. Recommended value is 55.
        Note:
            * Applicable only for 'file-based' vector stores.
        Types: int
 
    chunk_overlap:
        Optional Argument.
        Specifies the number of overlapping characters between two consecutive chunks
        to be used during the splitting of the input file.
        Note:
            * Applicable only for 'file-based' vector stores.
        Default Value: 150
        Types: int
 
    overwrite_object:
        Optional Argument.
        Specifies whether to overwrite the existing object with
        the same name in the database.
        Note:
            * Applicable only for 'file-based' vector stores.
        Types: bool
 
    ingest_params:
        Optional Argument.
        Specifies the parameters to be used for document ingestion for NIM.
        Notes:
            * Applicable only for NVIDIA NIM endpoint.
            * Applicable only for 'file-based' vector stores.
            * Refer to the IngestParams class for more details.
        Types: IngestParams
 
RETURNS:
    TeradataVectorStore instance.
 
RAISES:
    TeradataMLException.
 
EXAMPLES:
    # Example 1: Create a file-based vector store from all the PDF
    #            files in a directory by passing the directory 
    #            path in "documents" and 'amazon.titan-embed-text-v1' model
    #            in "embedding" as a string.
    # Initialize the required imports.
    >>> from langchain_teradata import TeradataVectorStore
 
    # Get the absolute path of the directory.
    >>> files = "<Enter/path/to/pdfs>"
    >>> vs_instance = TeradataVectorStore.from_documents(name="vs_example_1",
                                                         documents=files,
                                                         embedding="amazon.titan-embed-text-v1")
 
    # Example 2: Create a file-based vector store from
    #            a list of PDF files by passing the list of file names in
    #            "documents" and "embedding" as a TeradataAI object 
    #            of api_type 'aws' and model_name 'amazon.titan-embed-text-v1'.
    # 
    # Initialize the TeradataAI object using environment variables.
    >>> import os
    >>> os.environ["AWS_DEFAULT_REGION"] = "<Enter AWS Region>"
    >>> os.environ["AWS_ACCESS_KEY_ID"] = "<Enter AWS Access Key ID>"
    >>> os.environ["AWS_SECRET_ACCESS_KEY"] = "<Enter AWS Secret Key>"
    >>> os.environ["AWS_SESSION_TOKEN"] = "<Enter AWS Session key>"
    >>> llm_aws = TeradataAI(api_type="aws",
                             model_name="amazon.titan-embed-text-v2:0")
 
    # Create the 'file-based' vector store instance.
    >>> files = ["file1.pdf", "file2.pdf"]
    >>> vs_instance = TeradataVectorStore.from_documents(name="vs_example_2",
                                                         documents=files,
                                                         embedding=llm_aws)
 
    # Example 3: Create a content-based vector store by passing a list of 
    #            LangChain Document objects in "documents", Langchain 
    #            BedrockEmbeddings object in "embedding" and BedrockChatModel
    #            object in "chat_completion_model".
    # Initialize the required imports.
    >>> from langchain_aws import BedrockEmbeddings
    >>> from langchain_core.documents import Document
    >>> from langchain.chat_models import init_chat_model
 
    # Initialize the BedrockEmbeddings object, Document object and BedrockChatModel object.
    >>> llm_bedrock = BedrockEmbeddings(model_id="amazon.titan-embed-text-v1")
    >>> doc_lc = [Document(page_content="This is a test document.", id="doc1"),
    >>>           Document(page_content="This is another document.", id="doc2")]
    >>> model = init_chat_model("anthropic.claude-3-5-sonnet-20240620-v1:0", 
                                model_provider="bedrock_converse",
                                region_name="<Enter AWS Region>")
 
    # Create the 'content-based' vector store instance.
    >>> vs_instance = TeradataVectorStore.from_documents(name="vs_example_3",
                                                         documents=doc_lc,
                                                         embedding=llm_bedrock,
                                                         chat_completion_model=model)