diff --git a/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent/agent-bastian-response.png b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent/agent-bastian-response.png
new file mode 100644
index 0000000000..c019052277
Binary files /dev/null and b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent/agent-bastian-response.png differ
diff --git a/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent/agent-json-schema-field.png b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent/agent-json-schema-field.png
new file mode 100644
index 0000000000..6b66256eee
Binary files /dev/null and b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent/agent-json-schema-field.png differ
diff --git a/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-example.html b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-example.html
new file mode 100644
index 0000000000..0d4b9eddd5
--- /dev/null
+++ b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-example.html
@@ -0,0 +1,191 @@
+
+
+
+
+
+Invoice - #INV-001
+
+
+
+
+
+
+
+
+
+
Bill To:
+
John Doe
+ 789 Client St.
+ Los Angeles, CA 90001
+ john.doe@example.com
+
+
+
+
+
+
+ | Description |
+ Quantity |
+ Unit Price |
+ Total |
+
+
+
+
+ | Web Design Services |
+ 1 |
+ $1,200.00 |
+ $1,200.00 |
+
+
+ | SEO Optimization |
+ 5 hours |
+ $80.00 |
+ $400.00 |
+
+
+ | Domain Hosting (1 Year) |
+ 1 |
+ $150.00 |
+ $150.00 |
+
+
+
+
+
+
+
+ | Subtotal: |
+ $1,750.00 |
+
+
+ | Tax (10%): |
+ $175.00 |
+
+
+ | Total: |
+ $1,925.00 |
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-example.png b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-example.png
new file mode 100644
index 0000000000..3acfa58f9d
Binary files /dev/null and b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-example.png differ
diff --git a/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-marked.png b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-marked.png
new file mode 100644
index 0000000000..3535348768
Binary files /dev/null and b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-reference-marked.png differ
diff --git a/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-skills-and-tools-model.png b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-skills-and-tools-model.png
new file mode 100644
index 0000000000..bb1db45666
Binary files /dev/null and b/docs/assets/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition/ocr-skills-and-tools-model.png differ
diff --git a/docs/assets/user-guide/etendo-copilot/bundles/overview/copilot-feedback.png b/docs/assets/user-guide/etendo-copilot/bundles/overview/copilot-feedback.png
new file mode 100644
index 0000000000..ec3180aaee
Binary files /dev/null and b/docs/assets/user-guide/etendo-copilot/bundles/overview/copilot-feedback.png differ
diff --git a/docs/assets/user-guide/etendo-copilot/setup/agent-memory-window.png b/docs/assets/user-guide/etendo-copilot/setup/agent-memory-window.png
new file mode 100644
index 0000000000..6bdc762cde
Binary files /dev/null and b/docs/assets/user-guide/etendo-copilot/setup/agent-memory-window.png differ
diff --git a/docs/assets/user-guide/etendo-copilot/setup/assistant-window.png b/docs/assets/user-guide/etendo-copilot/setup/assistant-window.png
index 5f33f2d565..4a7230eae3 100644
Binary files a/docs/assets/user-guide/etendo-copilot/setup/assistant-window.png and b/docs/assets/user-guide/etendo-copilot/setup/assistant-window.png differ
diff --git a/docs/assets/user-guide/etendo-copilot/setup/knowledge-tab.png b/docs/assets/user-guide/etendo-copilot/setup/knowledge-tab.png
index 11715a0f99..db48752dae 100644
Binary files a/docs/assets/user-guide/etendo-copilot/setup/knowledge-tab.png and b/docs/assets/user-guide/etendo-copilot/setup/knowledge-tab.png differ
diff --git a/docs/assets/user-guide/etendo-copilot/setup/skill-tool-window.png b/docs/assets/user-guide/etendo-copilot/setup/skill-tool-window.png
index 9f6b01a0b3..cd1e1ec0fa 100644
Binary files a/docs/assets/user-guide/etendo-copilot/setup/skill-tool-window.png and b/docs/assets/user-guide/etendo-copilot/setup/skill-tool-window.png differ
diff --git a/docs/assets/user-guide/etendo-copilot/setup/skills-and-tools-tab.png b/docs/assets/user-guide/etendo-copilot/setup/skills-and-tools-tab.png
index a66cd4e968..6fb89db147 100644
Binary files a/docs/assets/user-guide/etendo-copilot/setup/skills-and-tools-tab.png and b/docs/assets/user-guide/etendo-copilot/setup/skills-and-tools-tab.png differ
diff --git a/docs/developer-guide/etendo-copilot/available-tools/ocr-tool.md b/docs/developer-guide/etendo-copilot/available-tools/ocr-tool.md
index 9bfeb318c4..ecebac9251 100644
--- a/docs/developer-guide/etendo-copilot/available-tools/ocr-tool.md
+++ b/docs/developer-guide/etendo-copilot/available-tools/ocr-tool.md
@@ -9,11 +9,19 @@ tags:
# Optical Character Recognition (OCR) Tool
-:octicons-package-16: Javapackage: `com.etendoerp.copilot.ocrtool`
+:octicons-package-16: Javapackage: `com.etendoerp.copilot.toolpack`
## Overview
-The Optical Character Recognition (OCR) Tool is a tool that recognizes text from images or pdfs. It can be used in Agents to extract information from images or pdfs that are uploaded to the chat.
+The Optical Character Recognition (OCR) Tool is an advanced tool that extracts structured data from images and PDF documents. It uses vision AI models to recognize and extract text, with support for automatic reference template matching, structured output schemas, and multi-provider configuration.
+
+**Key Features:**
+
+- **Automatic Reference Matching**: Searches for similar reference templates in the agent's vector database to guide extraction with visual markers
+- **Structured Output Schemas**: Supports predefined schemas (testdocument, etc.) and extensible custom schemas
+- **Multi-format Support**: JPEG, PNG, WebP, GIF, and multi-page PDF files
+- **Multi-provider Support**: Compatible with OpenAI (GPT-4o, GPT-5-mini) and Gemini models
+- **Advanced Configuration**: Per-agent model configuration, PDF quality control, and threshold filtering
!!!info
To be able to include this functionality, the Copilot Extensions Bundle must be installed. To do that, follow the instructions from the marketplace: [Copilot Extensions Bundle](https://marketplace.etendo.cloud/?#/product-details?module=82C5DA1B57884611ABA8F025619D4C05){target="\_blank"}. For more information about the available versions, core compatibility and new features, visit [Copilot Extensions - Release notes](../../../whats-new/release-notes/etendo-copilot/bundles/release-notes.md).
@@ -22,34 +30,192 @@ The Optical Character Recognition (OCR) Tool is a tool that recognizes text from
This tool automates the process of **text extraction from image-based files or PDFs**. This can be particularly useful for tasks such as document digitization, data extraction, and content analysis.
+The advanced features of the OCR Tool allow it to leverage reference templates and structured output schemas, enhancing the accuracy and consistency of the extracted data.
+
+When the tool is invoked, it processes the provided image or PDF file, searches for similar reference templates if available, and extracts the relevant information based on the specified question or instructions.
+
+**Threshold Filtering**: The tool can filter reference templates based on a configurable similarity threshold (configured in the `gradle.properties` file), like a minimum match score. This ensures that only closely matching references are used for extraction, improving accuracy. This threshold can be disabled if needed, allowing the tool to always use the best match found regardless of similarity level.
+
Using this tool consists of the following actions:
-- Receiving Parameters:
+- **Receiving Parameters**:
+
+ The tool receives an input object that contains the following parameters:
+
+ - **path** (required): The absolute or relative path of the image or PDF file to be processed. The file must exist in the local file system.
+
+ - **question** (required): A contextual question or instructions specifying the information to be extracted from the image. Be precise about the data fields needed (e.g., 'Extract invoice number, date, total amount, and vendor name'). Clear instructions improve extraction accuracy.
+
+ - **structured_output** (optional): Specify a schema name to use structured output format (e.g., 'testdocument'). Available schemas are loaded from `tools/schemas/` directory. When specified, the response will follow the predefined schema structure. Leave empty for unstructured JSON extraction.
+
+ - **scale** (optional): PDF render scale factor (e.g., 2.0 = ~200 DPI, 3.0 = ~300 DPI). Higher values yield better quality but larger size and slower processing. Default: 2.0.
+
+ - **disable_threshold_filter** (optional): Default: `false`. When `true`, ignore the configured similarity threshold and return the most similar reference found in the agent database (disables threshold filtering). In other words, the tool will always use the best match found, regardless of how different it might be from the original document.
+
+ - **force_structured_output_compat** (optional): Default: `false`. Modifies the communication method for structured output requests. By default it's false (uses native structured output system). When `true` (or automatically for models starting with 'gpt-5'), it changes how the structured input is sent to the LLM by bypassing the native structured-output wrapper and embedding the schema JSON directly into the system prompt. This ensures compatibility with older agents, specific model requirements, or for resolving compatibility issues. Use only when necessary.
+
+ !!! tip
+ To learn how to optimize the results of this tool, check the [How to Improve OCR Recognition](../how-to-guides/how-to-improve-ocr-recognition.md) guide.
+
+- **Obtaining the File**: The tool retrieves the file specified in the **path** parameter. It verifies the existence of the file and ensures it is in a supported format (`JPEG`, `JPG`, `PNG`, `WEBP`, `GIF`, `PDF`).
+
+- **PDF Conversion**: If the input file is a `PDF`, it is converted to an image format (`JPEG`) using the **pypdfium2** library. Each page of the `PDF` is rendered as a separate image.
+
+- **Image Conversion**: Other image formats are processed directly or converted to `JPEG` if necessary.
+
+- **Image Processing**: The image is processed using a Vision AI model (OpenAI GPT or Google Gemini, depending on configuration). This model interprets the text within the image and extracts the relevant information based on the provided **question**.
+
+- **Returning the Result**: The tool returns a JSON object containing the extracted information from the image or PDF.
+
+## Advanced Features
+
+### Automatic Reference Template Matching
+
+The OCR Tool includes an intelligent reference system that automatically searches for similar document templates in the agent's vector database. When a similar reference is found, it guides the extraction process by indicating which data fields to extract.
+
+**How it works:**
+
+1. When processing a document, the tool searches the agent's vector database (*ChromaDB*) for similar reference images.
+2. Reference images contain visual markers (red boxes) highlighting the relevant data fields to extract.
+3. The tool uses the reference as a template to prioritize and extract the same fields from the current document.
+4. This significantly improves extraction accuracy for documents with consistent layouts (invoices, receipts, forms).
+
+**Managing Reference Images:**
+
+- Upload reference images to the agent's knowledge base as you would normally do with any other document.
+- Reference images should have visual markers (typically red boxes) indicating the data fields
+- Each agent maintains its own vector database of reference templates
+- The similarity threshold can be controlled or disabled using the `disable_threshold_filter` parameter
+
+**When to use references:**
+
+- Processing invoices, receipts, or forms with consistent layouts.
+- When you need to extract specific fields repeatedly from similar documents.
+- To improve extraction accuracy by providing visual guidance.
+
+**Regulating Similarity Threshold:**
+
+The similarity threshold determines how closely a document must match a reference template to be used. This is controlled via a property in the `gradle.properties` file:
+
+- **Property**: `copilot.reference.similarity.threshold`
+- **Metric**: Uses L2 distance (lower values are stricter, higher values are more flexible).
+- **Recommended Range**: `0.15` to `0.30`.
+
+**Default Behavior (Out of the box):**
+
+By default, the tool is configured to be **highly permissive**:
+
+- If the `copilot.reference.similarity.threshold` property is **not set**, the tool does not apply any distance filtering.
+- It will automatically search for the most similar reference in the database and **always use the best match found**, regardless of how different it might be from the original document.
+- This ensures that if you have only one reference template, the tool will always try to use it.
+
+**How to specify or disable the threshold:**
+
+- **To specify a threshold**: Set the `copilot.reference.similarity.threshold` property in the `gradle.properties` file to a float value (e.g., `0.20`). This will prevent the tool from using references that are too different.
+- **To disable threshold filtering**: If you have a threshold configured but want to ignore it for a specific request, set the `disable_threshold_filter` parameter to `true` in the tool input. This will force the tool to use the most similar reference found, even if it's not a close match. This can be done by instructing the agent to disable the threshold in its prompt. The agent will then pass the parameter to the tool.
+
+### Structured Output Schemas
+
+The tool supports predefined schemas that enforce a specific output structure. This is useful when you need consistent data formats for downstream processing.
+
+**Available schemas:**
+
+Schemas are stored in the `tools/schemas/` directory and can be extended with custom schemas. In Copilot Extensions are included an example schema:
+
+- **testdocument**: Comprehensive test document schema with various field types (dates, amounts, line items, address, etc.)
+- Custom schemas can be added by creating new schema files in the schemas directory
+
+!!! tip
+ To learn how to create and use custom structured output schemas, check the [How to Improve OCR Recognition](../how-to-guides/how-to-improve-ocr-recognition.md#using-structured-output) guide.
- - The tool receives an input object that contains two keys:
+**Usage:**
- - **path**: The path of the image or PDF file to be processed.
- - **question**: A contextual question specifying the information to be extracted from the image. This is mandatory for precise results.
+```json
+{
+ "path": "/home/user/document.pdf",
+ "question": "Extract the information from this document",
+ "structured_output": "testdocument"
+}
+```
-- Obtaining the File:
+The tool will return data following the exact structure defined in the schema, ensuring consistency across all extractions.
- - The tool retrieves the file specified in the **path** parameter. It verifies the existence of the file and ensures it is in a supported format (JPEG, JPG, PNG, WEBP, GIF, PDF).
+**Compatibility mode:**
-- PDF Conversion:
+For older models or agents that don't support native structured output, use the `force_structured_output_compat` parameter. This changes how the structured input is sent to the LLM by embedding the schema JSON directly into the system prompt:
- - If the input file is a PDF, it is converted to an image format (JPEG) using the **pypdfium2** library. Each page of the PDF is rendered as a separate image.
+```json
+{
+ "path": "/home/user/document.pdf",
+ "question": "Extract the information from this document",
+ "structured_output": "testdocument",
+ "force_structured_output_compat": true
+}
+```
-- Image Conversion:
+### Multi-Provider Configuration
- - Other image formats are processed directly or converted to JPEG if necessary.
+The OCR Tool supports multiple AI providers with flexible configuration options:
-- Image Processing:
+**Supported Providers:**
- - The image is processed using a Vision model powered by GPT. This model interprets the text within the image and extracts the relevant information based on the provided **question**.
+- **OpenAI**: Models like `gpt-4o`, `gpt-5-mini` (default: `openai/gpt-5-mini`)
+- **Gemini**: Google's Gemini vision models
-- Returning the Result:
+**Configuration methods (in priority order):**
- - The tool returns a JSON object containing the extracted information from the image or PDF.
+1. **Global configuration (gradle.properties)**: Set `copilot.ocrtool.model` to configure the model globally for all agents. This property acts as a global override; if it is defined, the tool will use this model for all OCR requests, bypassing any per-agent configuration.
+
+ To configure it, add the property to your `gradle.properties` file specifying the **model name**:
+ ```properties
+ copilot.ocrtool.model="gpt-4o"
+ ```
+ The tool automatically infers the provider based on how the model name starts (e.g., models starting with `gemini` will use the Gemini provider, otherwise OpenAI is used).
+
+2. **Per-agent configuration**: Configure the model in the **Skills and Tools** tab of the Agent window using the **Model** field. The model must be specified using the format `provider/modelname` (e.g., `openai/gpt-4o`, `google/gemini-1.5-pro`). This allows different agents to use different models for the same tool.
+
+3. **Default**: If no configuration is provided, uses `openai/gpt-5-mini` with OpenAI provider
+
+**Provider detection:**
+
+The provider is automatically detected based on the model name:
+- Models starting with `gemini` use Gemini provider
+- Otherwise, the OpenAI provider is used by default
+
+### PDF Quality Control
+
+For `PDF` documents, you can control the rendering quality using the `scale` parameter:
+
+**Scale values:**
+
+- `2.0`: ~200 DPI (default) - Good balance between quality and performance
+- `3.0`: ~300 DPI - Higher quality, recommended for documents with small text
+- `4.0`: ~400 DPI - Maximum quality, slower processing and larger memory usage
+
+**Example:**
+
+```json
+{
+ "path": "/home/user/contract.pdf",
+ "question": "Extract all contract clauses",
+ "scale": 3.0
+}
+```
+
+**Considerations:**
+
+- Higher scale values produce better OCR accuracy for small or complex text
+- Increases processing time and memory consumption
+- Choose based on your document's text size and complexity
+
+### Performance Optimizations
+
+The tool includes several optimizations for better performance:
+
+- **In-memory processing**: Images and `PDFs` are converted to `Base64` directly in memory without disk I/O
+- **Efficient PDF rendering**: Uses *pypdfium2* for fast PDF-to-image conversion
+- **Automatic cleanup**: Temporary files are automatically removed after processing
+- **Parallel page processing**: Multi-page `PDFs` are processed efficiently
## Usage Example
@@ -57,7 +223,7 @@ Using this tool consists of the following actions:
### Requesting text recognition from an image/pdf
-Suppose you have an image at `/home/user/invoice.png` and you want to extract text related to an invoice information:
+Suppose you have an image at `/home/user/invoice.png` and you want to extract invoice information:
The following is an example image of an invoice:
@@ -198,6 +364,81 @@ The following is an example image of an invoice:
}
```
+### Advanced Usage: Structured Output with Reference
+
+Suppose you have uploaded a reference document image to your agent's vector database, and now you want to extract data with a predefined schema:
+
+- Use the tool with structured output:
+
+ - Input:
+
+ ```json
+ {
+ "path": "/home/user/document.pdf",
+ "question": "Extract the document information following the reference template",
+ "structured_output": "testdocument",
+ "scale": 3.0,
+ "disable_threshold_filter": false
+ }
+ ```
+
+ - The tool will:
+ 1. Search for a similar reference document in the vector database
+ 2. Use the reference's visual markers to guide extraction
+ 3. Render the PDF at 300 DPI for better quality
+ 4. Return data following the `testdocument` schema structure
+
+ - Output:
+
+ ```json
+ {
+ "document_id": "DOC-12345",
+ "document_number": "TX-2024-001",
+ "title": "Shipping Manifest",
+ "status": "pending",
+ "priority": "high",
+ "creation_date": "2024-12-16",
+ "owner": {
+ "name": "Jane Smith",
+ "email": "jane@example.com"
+ },
+ "line_items": [
+ {
+ "line_number": 1,
+ "item_code": "PROD-A",
+ "description": "Product A",
+ "quantity": 10,
+ "unit_price": 50.00,
+ "line_total": 500.00
+ }
+ ],
+ "is_active": true,
+ "version": 1
+ }
+ ```
+
+### Usage with Custom Model Configuration
+
+You can configure a specific model for OCR processing in multiple ways:
+
+**Option 1: Using the `gradle.properties` file (global configuration)**
+Open the `gradle.properties` file and set the following property to specify the model for all agents:
+```bash
+copilot.ocrtool.model="gpt-4o"
+```
+
+**Option 2: Using the Agent window (per-agent configuration)**
+
+1. Open the **Agent** window in Etendo Classic
+2. Go to the **Skills and Tools** tab
+3. Select the OCR Tool from the list
+4. In the **Model** field, specify the model using the format `provider/modelname`:
+ - For OpenAI: `openai/gpt-4o` or `openai/gpt-5-mini`
+ - For Gemini: `google/gemini-1.5-pro`
+5. If left empty, the tool will use the default model (`openai/gpt-5-mini`)
+
+This allows you to use different models for different agents depending on their specific needs (e.g., high accuracy for invoices vs. speed for receipts).
+
### Result Chaining
!!!note
diff --git a/docs/developer-guide/etendo-copilot/available-tools/task-creator-tool.md b/docs/developer-guide/etendo-copilot/available-tools/task-creator-tool.md
index 7e33ba7ec7..63c203e484 100644
--- a/docs/developer-guide/etendo-copilot/available-tools/task-creator-tool.md
+++ b/docs/developer-guide/etendo-copilot/available-tools/task-creator-tool.md
@@ -20,7 +20,7 @@ The **Task Creator Tool** automates the creation of tasks based on the content o
To be able to include this functionality, the Copilot Extensions Bundle must be installed. To do that, follow the instructions from the marketplace: [Copilot Extensions Bundle](https://marketplace.etendo.cloud/?#/product-details?module=82C5DA1B57884611ABA8F025619D4C05){target="\_blank"}. For more information about the available versions, core compatibility and new features, visit [Copilot Extensions - Release notes](../../../whats-new/release-notes/etendo-copilot/bundles/release-notes.md).
!!! tip
- To know when is the best time to use this tool or how there are executed this tasks, check the [How to create bulk tasks for Copilot](../how-to-guides/how-to-create-and-work-with-bulk-tasks-for-copilot.md) guide.
+ To know when is the best time to use this tool or how these tasks are executed, check the [How to create bulk tasks for Copilot](../how-to-guides/how-to-create-and-work-with-bulk-tasks-for-copilot.md) guide.
This tool provides the agent with:
@@ -52,9 +52,11 @@ The tool follows these main steps:
- `question`: Description or request that will be used as the task base. It is recommended that this be a singularized question. For example: "Process the product".
- `file_path`: Path to the input file (ZIP, CSV, XLS, or XLSX).
- `group_id`: Optional group ID. If not set, it uses the conversation ID.
+ - `groupby`: Optional list of column names to group rows by when processing CSV/XLS/XLSX files. If provided, rows that share the same values for these columns will be grouped together and sent as a single task. The task item will be a JSON array string containing all rows in the group. This is useful when multiple rows belong to the same entity (e.g., all order lines for the same order document number). Accepts a list or comma-separated string.
- `task_type_id`: Optional task type ID. If not provided, it uses the default "Copilot" task type (ID: `A83E397389DB42559B2D7719A442168F`).
- `status_id`: Optional status ID. If not provided, it uses the default "Pending" status (ID: `D0FCC72902F84486A890B70C1EB10C9C`).
- `agent_id`: Optional ID of the agent to process the task. If not provided, it uses the current main agent's ID.
+ - `preview`: Optional boolean. If true, returns the column names (or file contents for ZIP) instead of creating tasks.
- **File Extraction**
@@ -90,6 +92,8 @@ The tool follows these main steps:
## Usage Example
+### Example 1: Basic Usage
+
You have a `.csv` file with a list of customer feedback. Each row should become a separate task under the same group. You would input:
- `question`: Review feedback
@@ -101,7 +105,35 @@ You have a `.csv` file with a list of customer feedback. Each row should become
The Task Creator Tool will process each row and generate a task with the base question + the row's data. If a parameter is not set, the tool uses its default value.
-This helps you automate large-scale task generation in just one step, saving time and avoiding manual entry.
+### Example 2: Grouping by Column Values
+
+You have a `.xlsx` file with order lines where multiple rows belong to the same order document. Instead of creating one task per line, you want to create one task per order containing all its lines.
+
+**Sample Excel file:**
+
+| DocumentNo | Product | Quantity | Price |
+|------------|---------|----------|-------|
+| ORD-001 | Product A | 10 | 100 |
+| ORD-001 | Product B | 5 | 50 |
+| ORD-002 | Product C | 20 | 200 |
+| ORD-002 | Product D | 15 | 150 |
+
+**Input parameters:**
+
+ - `question`: Process order
+ - `file_path`: /path/to/orders.xlsx
+ - `group_id`: 123456789
+ - `groupby`: ["DocumentNo"]
+ - `agent_id`: (optional)
+
+**Result:**
+
+The tool will create **2 tasks** (one per order):
+
+1. **Task 1** with item data as a JSON array containing both lines for ORD-001
+2. **Task 2** with item data as a JSON array containing both lines for ORD-002
+
+This helps you automate large-scale task generation in just one step, saving time and avoiding manual entry, while maintaining the logical grouping of related data.
---
This work is licensed under :material-creative-commons: :fontawesome-brands-creative-commons-by: :fontawesome-brands-creative-commons-sa: [ CC BY-SA 2.5 ES](https://creativecommons.org/licenses/by-sa/2.5/es/){target="_blank"} by [Futit Services S.L](https://etendo.software){target="_blank"}.
diff --git a/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent.md b/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent.md
index 0b099dac4a..26b5900d86 100644
--- a/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent.md
+++ b/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-an-agent.md
@@ -99,9 +99,10 @@ After saving the agent, the system will automatically grant access to it. Open t
Currently, Copilot supports the following providers:
- **OpenAI**: This provider is the default one and is the most used. It is the most versatile and has the best performance in most cases.
+- **Google Gemini**: This provider is specialized in general tasks like OpenAI, but with better performance in some cases.
- **Anthropic**: This provider is specialized in code generation. It is the best option for code-related tasks.
- **Deepseek**: This provider is for general tasks like OpenAI, but cheaper.
-- **Ovider is for users that have their own models running in their own infrastructure. The support for this provider is in experimental phase. For more information visit, [How to Use and Run Self Hosted Models with Ollama](how-to-use-run-self-hosted-models-with-ollama.md) guide.llama (Self-hosted models)**: This pro
+- **Ollama (Self-hosted models)**: This provider is for users that have their own models running in their own infrastructure. The support for this provider is in experimental phase. For more information visit the [How to Use and Run Self Hosted Models with Ollama](how-to-use-run-self-hosted-models-with-ollama.md) guide.
### Default Model
The default model for Etendo Copilot is `gpt-4.1` from **OpenAI**. This model is selected automatically if the agent hasn't a specific model selected.
@@ -118,9 +119,7 @@ Etendo Copilot provides a Window where you can see the available models and thei
Models that support image inputs can work with images attached to the conversation. If the model does not support image inputs, it's possible to fix this by adding to the agent the `OCR Tool` that allows extracting text from images.
-This tool is available in the [Etendo Copilot ToolPack](../../../developer-guide/etendo-copilot/bundles/overview.md#etendo-copilot-toolpack) module.
-
-The OCR Tool is a tool that allows extracting text and information from images. Maybe it's necessary to **explain** to the agent how to use it in the prompt.
+This tool is available in the [Etendo Copilot ToolPack](../../../developer-guide/etendo-copilot/bundles/overview.md#etendo-copilot-toolpack) module. You may need to **explain** to the agent how to use it in the prompt.
@@ -140,20 +139,130 @@ The most crucial is to determine:
|**Remote File** | It is highly recommended when the file can change and the latest version can be accessed from the same url. For example, a file in a repository on GitHub. | File URL |
|**HQL Query** | It is used when you want the agent to be able to read information from a table or from the result of a database query. For example, a list of Business Partners or orders. | HQL Query |
|**Text** | When the information is static and can be written directly in the window. | The text itself |
-|**OpenAPI Flow Specification** | Use when the knowledge base file is the OpenAPI Specification of a Etendo Classic Flow. See [How to allow Copilot to interact with Etendo Classic](#how-to-allow-copilot-to-interact-with-etendo-classic) for more information. | Select the flow in the selector|
+|**OpenAPI Flow Specification** | Use when the knowledge base file is the OpenAPI Specification of a Etendo Flow. See [How to allow Copilot to interact with Etendo](#how-to-allow-copilot-to-interact-with-etendo-classic) for more information. | Select the flow in the selector|
|**Code Index** | When the agent needs to know **Locally** stored code. | Specify the paths of the folders |
!!! info
More information about this window can be found in the [Knowledge Base File Window](../../../user-guide/etendo-copilot/setup-and-usage.md#knowledge-base-file-window) article.
-### Advanced settings
+!!!info "Image Files in Knowledge Base"
+ When files are indexed in an agent's knowledge base, **image files are handled separately** from text documents:
+
+ - **Text documents** are indexed in the main vector database for semantic search
+ - **Image files** (PNG, JPG, JPEG) are indexed in a **separate image database** for visual similarity search
+ - This image database is used by tools like the [OCR Tool](../available-tools/ocr-tool.md) to find reference templates with visual markers
+ - The OCR Tool automatically searches this database to find similar reference images that guide data extraction
+ - Each agent maintains its own image database, independent from its text knowledge base
+
+### Advanced settings
In the Knowledge Base File window, there is an advanced settings section that allows you to configure the following options in the splitting algorithm of the content of the file:

- **Skip Splitting**: Retrieves the entire document as one chunk, which is useful for small files.
-- **Max. Chunk Size**: This option allows to set the maximum size (tokens) of the chunks that will be created when the content is split. This is useful to avoid very large chunks that can cause performance issues. Depending on the file types, the splitting algorithm checks for **separators** to split the content semantically. For example, in markdown files, the splitting is done by headers, so each chunk will contain the content of a header and its subheaders. Or in the case of Java files, the splitting is done by classes, so each chunk will contain the content of a class and its methods. When the chunk size is reached, the content is split into a new chunk in the next separator found. This is useful to avoid very large chunks that can cause problems with the token limit of the model.
-- **Chunk Overlap**: This option allows to set the overlap between chunks. This is useful to avoid losing information when the content is split into chunks. The overlap is the number of tokens that are repeated in each chunk. For example, if the chunk size is 100 and the overlap is 10, each chunk will contain 90 unique tokens and 10 repeated tokens from the previous chunk. This is useful to avoid losing information when the content is split into chunks. Can be 0 if you don't want to have overlap between chunks.
+- **Max. Chunk Size**: This option allows you to set the maximum size (in tokens) of the chunks that will be created when the content is split. This is useful to avoid very large chunks that can cause performance issues. Depending on the file types, the splitting algorithm checks for **separators** to split the content semantically. For example, in markdown files, the splitting is done by headers, so each chunk will contain the content of a header and its subheaders. Or in the case of Java files, the splitting is done by classes, so each chunk will contain the content of a class and its methods. When the chunk size is reached, the content is split into a new chunk in the next separator found. This is useful to avoid very large chunks that can cause problems with the token limit of the model.
+- **Chunk Overlap**: This option allows you to set the overlap between chunks to avoid losing information when splitting. The overlap is the number of tokens repeated in each chunk. For example, if the chunk size is 100 and the overlap is 10, each chunk will contain 90 unique tokens and 10 repeated tokens from the previous chunk. Can be 0 if you don't want overlap between chunks.
+
+
+## Add Structured Outputs (JSON Schema)
+
+Starting in the latest release, the Agent configuration window includes a new field in the **Advanced Settings** section named **`JSON Schema for Structured Outputs`**. This field accepts a JSON Schema object which the agent will use to validate and format its responses.
+
+How to use the field:
+
+1. Open the Agent window and switch to the **Advanced Settings** tab.
+2. Locate the **`JSON Schema for Structured Outputs`** field.
+3. Paste a valid JSON Schema object into the field.
+4. Save the agent configuration and test with sample prompts.
+
+Important details:
+
+- The content must be valid JSON representing a JSON Schema. Use a JSON validator if unsure.
+- The agent will validate and attempt to format its response to match the schema.
+- If the agent fails to fully satisfy the schema, it will return a structured error describing validation issues.
+
+Example schema (Person record):
+
+```json
+{
+ "type": "object",
+ "properties": {
+ "name": {"type": "string"},
+ "email": {"type": "string", "format": "email"},
+ "department": {"type": "string"}
+ },
+ "required": ["name", "email"]
+}
+```
+
+Use cases:
+
+- Returning structured customer data for downstream processing.
+- Ensuring task extraction follows a predictable schema for integrations.
+- Producing event objects ready for ingestion by downstream systems.
+
+### Example: Multilingual Content (Bastian agent)
+
+Below is an example JSON Schema used in the `Bastian` agent (the Etendo wiki indexed) to return multilingual content and suggested questions.
+
+```json
+{
+ "$schema": "https://json-schema.org/draft/2020-12/schema",
+ "$id": "https://example.com/content-object.schema.json",
+ "title": "Multilingual Content Object",
+ "description": "An object containing content in English, its Spanish translation, and related suggested questions.",
+ "type": "object",
+ "properties": {
+ "content_en": {
+ "type": "string",
+ "description": "The primary content of the object in the English language."
+ },
+ "suggested_questions": {
+ "type": "array",
+ "description": "A list of suggested questions related to the main content.",
+ "items": {
+ "type": "string",
+ "description": "A single suggested question string."
+ },
+ "minItems": 1,
+ "uniqueItems": true
+ }
+ },
+ "required": [
+ "content_en",
+ "suggested_questions"
+ ],
+ "additionalProperties": false
+}
+```
+
+#### Explanation of the structure:
+
+- **$schema / title / description**: Metadata that documents the schema and the draft used.
+- **type: object**: The response must be a JSON object.
+- **properties**:
+ - `content_en` (string): Primary English content.
+ - `suggested_questions` (array[string]): One or more suggested question strings related to the content. `minItems: 1` enforces at least one suggestion; `uniqueItems: true` avoids duplicates.
+- **required**: Ensures `content_en` and `suggested_questions` are always present.
+- **additionalProperties: false**: Prevents extra fields beyond those declared; helps keep output strictly typed.
+
+Why this schema is useful:
+
+- Ensures the agent always returns the English text plus actionable suggested questions.
+- Guarantees programmatic parsing without fragile text extraction.
+
+Example usage (Bastian):
+
+- Configure the `Bastian` agent's Advanced Settings `JSON Schema for Structured Outputs` with the schema above.
+- Ask the agent a question like: `Summarize the Etendo installation steps for end users and suggest follow-up questions.`
+- The agent will respond with a JSON object matching the schema.
+
+JSON Schema field configured in the Agent Advanced Settings:
+
+
+
+Question asked to `Bastian` and structured JSON response received:
+
### Knowledge Base File Behavior
@@ -169,11 +278,11 @@ In the Knowledge Base File window, there is an advanced settings section that al
!!! tip
- **Remember the Synchronization**: After adding/modifying/deleting a knowledge base file from an Agent, its necessary to synchronize the agent to apply the changes. This not only regenerates/reloads the Knowledge Base File but also updates the Agent with the latest changes.
- - **Splitting**: We the indexation in the knowledge base file is done, the content is splitted in chunks depending of the type of the file. For example, if the file is a markdown file, the content is splitted in chunks by the headers. If the files are not large, its possible to mark as `Skip Splitting` in the knowledge base file configuration. This will avoid the splitting of the content in chunks. This causes that the content of the documents is retrieved as a single chunk, which can be useful in some cases.
+ - **Splitting**: When the indexation in the knowledge base file is done, the content is split into chunks depending on the type of the file. For example, if the file is a markdown file, the content is splitted in chunks by the headers. If the files are not large, its possible to mark as `Skip Splitting` in the knowledge base file configuration. This will avoid the splitting of the content in chunks. This causes that the content of the documents is retrieved as a single chunk, which can be useful in some cases.
### Add a Knowledge Base Example
-We got the example of the default Copilot agent `Bastian` that has a knowledge base file based in the Etendo Documentation from it GitHub repository. Copilot supports `.zip` format for the knowledge base file behavior, automatically extracting it and indexing the files inside.
+The default Copilot agent `Bastian` has a knowledge base file based on the Etendo Documentation from its GitHub repository. Copilot supports `.zip` format for the knowledge base file behavior, automatically extracting it and indexing the files inside.
In this case, the `ZIP` file contains the Etendo Documentation in markdown format. The agent has the knowledge base file configured as `Remote File` and the behavior as `Add to the agent as Knowledge Base`. The agent has the following configuration:
- Setting the Knowledge Base File:
@@ -234,25 +343,25 @@ To add a tool to an agent, follow these steps:
For example, we will add a tool to the agent `Task Definition Agent` to allow to write a file with the task definition. The tool will be the `Write File Tool` that allows to write a file with the content provided.

-After adding the tool, the agent will have the tool available to use. The agent can use the tool to write a file with the task definition. The agent will use the tool to write the file with the task definition.
+After adding the tool, the agent will have the tool available to use. The agent can use the tool to write a file with the task definition.

We can check the created file:

-## How to allow Copilot to Interact with an API or Etendo Classic?
-The most powerful and useful feature of Etendo Copilot is the ability to interact with APIs (including Etendo Classi)c. Currently the paradigm of AI agents is to automate and/or reuse what is **already done**. In other words, the utility arises from the fact that AI agents can use all the business logic that is already available.
+## How to allow Copilot to Interact with an API or Etendo?
+The most powerful and useful feature of Etendo Copilot is the ability to interact with APIs (including Etendo API). Currently the paradigm of AI agents is to automate and/or reuse what is **already done**. In other words, the utility arises from the fact that AI agents can use all the business logic that is already available.
### External API
The most usual way is based on a combination of an OpenAPI Specification and a tool that allows to make requests to that API. To do this, the following steps are needed:
- **Add the OpenAPI Specification**: The OpenAPI Specification is a standard way to describe an API. This specification is added as a Knowledge Base File. And configure it as `[Agent] Append the file content to the prompt`. This will allow the agent to know the endpoints and methods of the API.
- **Add the API Call Tool**: The API Call Tool is a tool that allows to make requests to an API. This tool is added as a tool in the agent. The agent can use this tool to make requests to the API.
-### Etendo Classic
-For Etendo Classic, the process is a bit different. The main difference is that we can take advantage of the OpenAPI Specification automatically generated by the `Flows`, where we can define a set of endpoints to which we want to give access to our agent.
-To know more about how to create a flow in Etendo Classic, check the [How to Document an Endpoint with OpenAPI](../../etendo-classic/how-to-guides/how-to-document-an-endpoint-with-openapi.md) guide.
-The steps to allow an agent to interact with Etendo Classic are:
+### Etendo
+For Etendo, the process is a bit different. The main difference is that we can take advantage of the OpenAPI Specification automatically generated by the `Flows`, where we can define a set of endpoints to which we want to give access to our agent.
+To know more about how to create a flow in Etendo, check the [How to Document an Endpoint with OpenAPI](../../etendo-classic/how-to-guides/how-to-document-an-endpoint-with-openapi.md) guide.
+The steps to allow an agent to interact with Etendo are:
- **Add the OpenAPI Specification**: This specification is added as a Knowledge Base File of type `OpenAPI Flow Specification`. When this type is selected, a selector with the available flows is shown, to select the flow that we want to use. The behavior of this file can be `[Agent] Append the file content to the prompt`. This will allow the agent to know the endpoints and methods of the API.
- **Add the API Call Tool**: The API Call Tool is a tool that allows to make requests to an API. This tool is added as a tool in the agent. The agent can use this tool to make requests to the API.
@@ -266,12 +375,12 @@ When the OpenAPI Specification is added as a Knowledge Base File of type `OpenAP
### Example of Copilot Interaction with Etendo
-For example, we will create an agent to create Products in Etendo Classic, using an already defined flow with the endpoints needed to create Products, Product Categories and Prices.
+For example, we will create an agent to create Products in Etendo, using an already defined flow with the endpoints needed to create Products, Product Categories and Prices.
1. First, we will create a new Knowledge Base File of type `OpenAPI Flow Specification` and select the flow `Product Flow`.
!!!info
- Be sure to use the agent that has the necessary permissions to interact with the data. In this case, the agent must have the necessary permissions to create products, categories and prices in Etendo Classic.
+ Be sure to use the agent that has the necessary permissions to interact with the data. In this case, the agent must have the necessary permissions to create products, categories and prices in Etendo.

diff --git a/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-and-work-with-bulk-tasks-for-copilot.md b/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-and-work-with-bulk-tasks-for-copilot.md
index a0fa645a1d..d4d900efdc 100644
--- a/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-and-work-with-bulk-tasks-for-copilot.md
+++ b/docs/developer-guide/etendo-copilot/how-to-guides/how-to-create-and-work-with-bulk-tasks-for-copilot.md
@@ -10,7 +10,7 @@ tags:
## Overview
-This article explains how to create and work with bulk tasks for Copilot. This is useful when you want to create multiple tasks at once and execute in background with Copilot
+This article explains how to create and work with bulk tasks for Copilot. This is useful when you want to create multiple tasks at once and execute them in the background with Copilot.
### Concept and Use Cases
When you need to make use of an AI agent to perform tasks with a high volume of iterations, it will be limited in the amount it can handle and its speed. Then the concept of Bulk tasks is born, which consists of storing requests in a window of the `Tasks` module.
@@ -23,20 +23,25 @@ These requests can be executed manually or be processed in a background process
## Add Copilot Tasks
The Etendo Copilot module includes:
-- **Add Copilot Task Button**: This button in the `Tasks` window allows you sent a CSV/XLSX/ZIP file to create bulk tasks.
+- **Add Copilot Task Button**: This button in the `Tasks` window allows you to send a CSV/XLSX/ZIP file to create bulk tasks.
- **Bulk Task Creator**: This agent is configured with the [Task Creator Tool](../available-tools/task-creator-tool.md) to create bulk tasks based on a zip file or a CSV/XLSX file. This agent can be added to a supervisor agent to chain with other agents.
In both options, the requirements are the same:
- **CSV/XLSX/ZIP file**: The file that contains the data to be processed.
-- **Question**: The description or request that will be used as the task base. For the case of the agent `Bulk Task Creator`, it will be encharge to reform the task to convert to singular tasks. But for the other options, it is necessary to provide the task base in singular form.
-- **Execution Group**: Optional group. If not set, it uses the conversation ID. Its use to identify the tasks that belong to the same group.
-- **Task Type**: Optional task type ID. If not set, it auto-creates one named "Copilot". Its use to identify the type of task.
-- **Status**: Optional status ID. Defaults to "Pending". Its use to identify the status of the task. After the task is processed, the status will be updated to `Completed`.
+- **Question**: The description or request that will be used as the task base. For the case of the agent `Bulk Task Creator`, it will be in charge of reforming the task to convert it to singular tasks. But for the other options, it is necessary to provide the task base in singular form.
+- **Execution Group**: Optional group. If not set, it uses the conversation ID. It is used to identify the tasks that belong to the same group.
+- **Task Type**: Optional task type ID. If not set, it auto-creates one named "Copilot". It is used to identify the type of task.
+- **Status**: Optional status ID. Defaults to "Pending". It is used to identify the status of the task. After the task is processed, the status will be updated to `Completed`.
- **Agent**: The agent that will process the tasks. If the agent is not specified, will be selected the agent that used the tool. For the case of the agent `Bulk task creator`, it will be selected the supervisor agent that contains it.
+- **Groupby**: Optional parameter. A list of column names (or a comma-separated string) to group rows by their values. When specified, rows with the same values in these columns will be grouped into a single task, with the data passed as a JSON array. This is useful for processing related records together (e.g., all order lines for the same document). If not specified, each row becomes a separate task.
+- **Preview**: Optional boolean parameter. When set to `true`, the tool will return a preview of how the tasks would be created without actually creating them. This is useful for validating the task structure before creating a large number of tasks.
!!! info
- When the tasks are created, the tasks will be created one per row in the case of CSV/XLSX files, one per file in the case of ZIP files, and one per file in the case of other files. The task request will be the following format: ```BASE_TASK - [FILE_NAME/ROW DATA]```
+ When the tasks are created, the tasks will be created one per row in the case of CSV/XLSX files (unless `groupby` is used), one per file in the case of ZIP files, and one per file in the case of other files. The task request will be in the following format:
+
+ - **Without groupby**: `BASE_TASK - [FILE_NAME/ROW DATA]`
+ - **With groupby**: `BASE_TASK - [JSON array with grouped rows]`
### Example
For example, if you have a CSV file with the following data:
@@ -73,7 +78,7 @@ And the objective is to insert these products in Etendo. For this example we wil
1. Go to the agent that contains the `Task Creator Tool`.
2. Open a conversation with the agent.
-3. Attach in the conversation the CSV file with the products data.
+3. Attach to the conversation the CSV file with the products data.
4. Send some request like:
``` text
@@ -92,11 +97,11 @@ And the objective is to insert these products in Etendo. For this example we wil

#### Using the `Bulk Task Creator` agent
-This agent know to use the `Task Creator Tool` strategically, converting the request in singular tasks. The steps to create the bulk tasks are:
+This agent knows how to use the `Task Creator Tool` strategically, converting the request into singular tasks. The steps to create the bulk tasks are:
1. Add the `Bulk Task Creator` agent to a supervisor agent. In this case, we will use the `Data Initialization Supervisor` agent that contains the `Product Generator` agent.
2. Open a conversation with the supervisor agent.
-3. Attach in the conversation the CSV file with the products data.
+3. Attach to the conversation the CSV file with the products data.
4. Send some request like:
``` text
@@ -107,6 +112,62 @@ This agent know to use the `Task Creator Tool` strategically, converting the req
5. The supervisor agent will delegate the task to the `Bulk Task Creator` agent that will create the tasks based on the file data and then process it.
+### Advanced Options: Grouping Related Data
+
+When working with data that has relationships between rows (such as order lines belonging to the same order, or invoice items for the same invoice), you can use the `groupby` parameter to create one task per group instead of one task per row.
+
+#### When to Use Groupby
+
+Use the `groupby` parameter when:
+
+- **Related records should be processed together**: For example, all lines of a sales order should be processed as a single order creation task.
+- **Context from multiple rows is needed**: When the AI agent needs to see all related data at once to make better decisions.
+- **Reducing task overhead**: Instead of creating hundreds of individual tasks, you can create fewer tasks with grouped data.
+
+#### Groupby Example: Order Lines
+
+Suppose you have an Excel file with order lines data:
+
+```csv title="order_lines.csv"
+DocumentNo,Product,Quantity,Price
+ORD-001,Laptop,2,1200
+ORD-001,Mouse,2,25
+ORD-002,Keyboard,5,80
+ORD-002,Monitor,3,350
+```
+
+Without groupby, this would create **4 separate tasks** (one per row). With groupby, you can group by `DocumentNo` to create **2 tasks** (one per order).
+
+**Using the Task Creator Tool with groupby:**
+
+```text
+Create bulk tasks for the attached file.
+- Question: "Create a sales order with these lines:"
+- Group by: DocumentNo
+- The group id is 'Order Load December 2025'
+- The agent ID is 'A9E0861E88B1460A98CAF55DCB2BEE82'
+```
+
+This will create 2 tasks:
+
+1. **Task 1**: `Create a sales order with these lines: - [{"DocumentNo": "ORD-001", "Product": "Laptop", "Quantity": 2, "Price": 1200}, {"DocumentNo": "ORD-001", "Product": "Mouse", "Quantity": 2, "Price": 25}]`
+2. **Task 2**: `Create a sales order with these lines: - [{"DocumentNo": "ORD-002", "Product": "Keyboard", "Quantity": 5, "Price": 80}, {"DocumentNo": "ORD-002", "Product": "Monitor", "Quantity": 3, "Price": 350}]`
+
+Each task now contains a JSON array with all the lines for that specific order, allowing the agent to process the complete order in a single execution.
+
+#### Preview Mode
+
+Before creating a large number of tasks, you can use the `preview` parameter to validate the task structure:
+
+```text
+Create bulk tasks for the attached file in preview mode.
+- Question: "Create product with this data:"
+- Preview: true
+- The agent ID is 'A9E0861E88B1460A98CAF55DCB2BEE82'
+```
+
+The tool will return a preview showing how many tasks would be created and their structure, without actually creating them in the database.
+
## How to Process Copilot tasks
The tasks can be processed in two ways:
diff --git a/docs/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition.md b/docs/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition.md
new file mode 100644
index 0000000000..5aa751a1d9
--- /dev/null
+++ b/docs/developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition.md
@@ -0,0 +1,119 @@
+---
+title: How to Improve OCR Recognition
+tags:
+ - Copilot
+ - OCR
+ - Purchase Invoice Expert
+ - Best Practices
+---
+
+# How to Improve OCR Recognition
+
+In simple terms, the OCR tool makes a call to a model to detect information from a document or image. However, it is not infallible. There are different methods to increase its reliability and efficiency.
+
+
+
+
+The tool follows these steps:
+
+1. **Convert to Image**: If the input is a PDF, it is converted into an image. If it is already an image, it proceeds directly.
+2. **Reference Search**: It searches for reference images in the agent's knowledge base (if any are configured).
+3. **LLM Call**: It calls an LLM model using a prompt (the `question` parameter of the OCR tool), which specifies exactly what the model should recognize and extract.
+
+## Using a specific model
+By default, the tool uses a model configured in the `gradle.properties` file. However, it is highly recommended to configure a specific model for the agent that will use the tool to ensure the best performance for visual tasks.
+
+In the **Agent** window, under the **Skills and Tools** tab, you can specify a **Model** for the OCR tool. This field allows you to override the default model for that specific tool in that agent.
+
+
+
+- **Format**: The model must be specified using the format `provider/modelname` (e.g., `openai/gpt-5-mini`).
+- **Recommendation**: Use models with strong vision capabilities. New models with improved recognition accuracy and faster processing speeds are released frequently, so it is advisable to stay updated with the latest available versions.
+
+!!! info
+ The **Model** field only appears in the **Skills and Tools** tab if the tool has the **Use Model** checkbox enabled in the **Skill/Tool** window.
+
+## Using a fixed prompt
+In general, when the agent calls the OCR tool, it sends a `question` (prompt) about what it needs to recognize. The agent constructs this question dynamically based on the current context and needs. However, generating this question every time can lead to inconsistencies between executions.
+
+To avoid this, you can "fix" the prompt for the OCR tool by instructing the agent exactly what question it should use.
+
+- **How to do it**: In the agent's prompt, specify the exact instructions for the OCR tool. For example: *"When calling the OCR tool, use this question: 'Read the product name, quantities, tax used, and total amount of the receipt'"*.
+- **Benefit**: By fixing the question, you ensure that the OCR tool always receives the same instructions, making the extraction process more predictable. Furthermore, any improvements made to this fixed question will positively impact all executions.
+
+## Increasing DPI
+When the input file is a PDF, the OCR tool must render it into an image before processing. This conversion quality is controlled by a parameter called `scale`.
+
+- **Default Value**: The default scale is `3.0`, which corresponds to approximately **300 DPI**.
+- **How to Increase it**: To improve recognition of very small text or extremely complex layouts, you can increase the rendering quality by instructing the agent in its prompt. For example, you can add: *"Use a scale of 4.0 when using OCRTool"* to achieve **400 DPI**.
+
+Higher scale values yield better quality and more accurate extraction, but they also result in larger image sizes and slower processing times.
+
+## Using Structured Output
+The OCRTool supports **Structured Output**, which allows you to define a specific JSON schema that the model must follow when extracting information. This ensures that the output is always consistent and easy to process by other tools or systems.
+
+To use structured output, you must:
+
+1. **Define the Schema**: Create a Python file in the `tools/schemas/` directory located at the root of your Copilot module (e.g., `modules/my_module/tools/schemas/mail.py`). The file name will be the schema name (e.g., `mail.py`).
+2. **Implement the Pydantic Model**: Inside the file, define a class that inherits from `pydantic.BaseModel`. The class name must follow the pattern `Schema` (e.g., `class MailSchema(BaseModel):`).
+3. **Invoke the Tool**: Specify the schema name in the `structured_output` parameter when calling the OCRTool.
+
+### Example: Creating a "Mail" Schema
+
+Suppose you want to extract information from scanned emails or letters. You can create a schema named `Mail`:
+
+**File**: `tools/schemas/mail.py`
+
+```python
+from typing import List, Optional
+from pydantic import BaseModel, Field
+
+# Internal model for nested data
+class MailAttachment(BaseModel):
+ filename: str = Field(description="Name of the attachment file")
+ size: Optional[int] = Field(description="Size of the file in bytes")
+
+# Main schema class - Must end with 'Schema'
+class MailSchema(BaseModel):
+ """Schema for extracting email/mail information."""
+ sender: str = Field(description="The person or entity who sent the mail")
+ recipient: str = Field(description="The person or entity who received the mail")
+ subject: Optional[str] = Field(description="The subject line of the mail")
+ body: str = Field(description="The main content or body of the mail")
+ date: Optional[str] = Field(description="The date when the mail was sent")
+ attachments: List[MailAttachment] = Field(default_factory=list, description="List of attachments mentioned")
+```
+
+Once defined, you can instruct the agent to use this schema:
+
+- **Agent Prompt**: *"When you extract information from an email image, use the OCRTool with the 'Mail' structured output."*
+
+The tool dynamically loads the schema from the file. It converts the `structured_output` value (e.g., `Mail`) to lowercase to find the file (`mail.py`) and then looks for the class `MailSchema`.
+
+The resulting extraction will follow the exact structure of your `MailSchema`, making it highly reliable for automated workflows.
+
+## Adding reference data
+The OCR tool includes an intelligent reference system that automatically searches for similar document templates in the agent's knowledge base. When a similar reference is found, it guides the extraction process by indicating which data fields to extract using visual markers.
+
+To add reference data:
+
+1. **Prepare the Reference**: Take a high-quality image of the document type (e.g., a specific supplier's invoice). If possible, highlight the key fields that need to be extracted (e.g., total amount, tax, date) using colored boxes or hints.
+2. **Upload to Knowledge Base**: In the **Knowledge Base File** window, upload the image.
+3. **Index the Image**: Ensure the file is synchronized with the agent. The system will automatically index it in a separate image database designed for visual similarity search.
+4. **Automatic Matching**: When the agent processes a new document, the OCR tool will automatically search for the most similar reference in this database to improve extraction accuracy.
+
+**Example: Improving Business Partner detection in Purchase Invoice Expert**
+
+If you want to improve the detection of the Business Partner (BP) for a specific invoice format, you can add a reference image of that invoice to the agent's Knowledge Base. In this image, you should mark the sections where the BP information is located. This guides the OCR tool to look in those specific areas when it encounters a similar invoice layout, significantly increasing the reliability of the extraction.
+
+
+
+Marked fields in the reference image to guide OCR extraction.
+
+
+
+!!! tip
+ Using reference templates is especially useful for documents with non-standard layouts where the LLM might otherwise struggle to locate specific fields.
+
+!!! tip
+ If you don't want to use reference images for a specific execution, you can disable this feature by indicating in the agent's prompt: *"When using OCRTool, do not use reference images from the knowledge base."*. The agent will disable the reference search for that execution through a parameter of the OCR tool.
diff --git a/docs/user-guide/etendo-copilot/bundles/overview.md b/docs/user-guide/etendo-copilot/bundles/overview.md
index dfa77404d7..bf602218e9 100644
--- a/docs/user-guide/etendo-copilot/bundles/overview.md
+++ b/docs/user-guide/etendo-copilot/bundles/overview.md
@@ -87,6 +87,15 @@ This supervisor has the following agents:
- **Purchase Invoice Expert**: Agent expert in managing purchase invoices for Etendo. It manages the entire invoice creation process, extracts and validates the invoice header and lines. Finally, it invokes APIs to insert data and provides final validation.
+ When this agent processes an invoice, it performs a validation of the total amount. To provide feedback to the user, a **Copilot** section is added to the **Purchase Invoice** window with a field called **Copilot Feedback**. The agent updates this field with the result of the validation:
+ 
+
+ - **Ok**: The invoice total matches the extracted data and is ready for processing.
+ - **Requires manual review**: There is a mismatch or an issue that requires a human user to verify the invoice details.
+
+ !!! tip
+ To improve the accuracy of data extraction, check the [How to Improve OCR Recognition](../../../developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition.md) guide.
+
##### Order Expert
@@ -110,7 +119,7 @@ This agent reads a zip file and returns the paths of the files inside the zip. I
:octicons-package-16: Javapackage: `com.etendoerp.copilot.toolpack`
-The **Etendo Copilot Toolpack** is a collection of tools that help to **developers** to add functionalities to agents, such request to an API, send an email, read a file, write a file, and more.
+The **Etendo Copilot Toolpack** is a collection of tools that help **developers** add functionalities to agents, such request to an API, send an email, read a file, write a file, and more.
!!! info
For more information, visit the [Toolpack - Developer Guide](../../../developer-guide/etendo-copilot/bundles/overview.md#etendo-copilot-toolpack), where you will find a detailed list of the available tools, instructions on how to use them, and a guide for developing new tools.
diff --git a/docs/user-guide/etendo-copilot/setup-and-usage.md b/docs/user-guide/etendo-copilot/setup-and-usage.md
index 02ce509a0f..63ec2f2be4 100644
--- a/docs/user-guide/etendo-copilot/setup-and-usage.md
+++ b/docs/user-guide/etendo-copilot/setup-and-usage.md
@@ -79,6 +79,8 @@ Fields to note:
- **Search Result Qty.**: This option allows you to set the number of search results in the knowledge base on which the agent will base its response. The default value is 4, but it can be changed to any value. This value is useful when the agent has a large knowledge base, and you want to increase/decrease the number of results returned by the agent.
- **Temperature**: This controls randomness, lowering results in less random completions. As the temperature approaches zero, the model will become deterministic and repetitive.
+- **JSON Schema for Structured Outputs**: When configured, the agent will attempt to return responses that conform to the provided schema described in the JSON Schema format. This is useful for ensuring that the agent's outputs are structured and can be easily parsed or processed by other systems.
+
### Buttons
@@ -88,7 +90,7 @@ Fields to note:
- **Refresh Preview**: Show only when agent type is **Langraph**, allowing the user to refresh the Graph Preview when changes to the team members are introduced.
-- **Check hosts**: This button check the configuration of Etendo and Copilot, to ensure that de comunication between them is correct. In case of any error, a message will be shown.
+- **Check hosts**: This button checks the configuration of Etendo and Copilot, to ensure that the communication between them is correct. In case of any error, a message will be shown.
- **Clone**: The navbar clone button allows the cloning of agents, making a copy of both all header fields and related records in the tabs. When a agent is cloned in, the name `Copy of` is added.
@@ -117,6 +119,7 @@ Fields to note:
- **Active**: checkbox to activate the knowledge base file.
- **Type**: read-only field showing the type of file selected in the [Knowledge Base File window](#knowledge-base-file-window).
+- **Module**: Module in which this knowledge base file configuration will be exported. This field is only available with the `System Administrator` role.
- **Alias** In case you select behaviour, `[Agent] Append the file content to the prompt`, by default it adds the file content dynamically to the end of the prompt, the alias can be used to replace the file content inside the prompt, using the wildcard @@, with the alias you define in this field.
### Skills and Tools Tab
@@ -127,8 +130,10 @@ In this tab, you can define the tools to be used by the agent.
Fields to note:
-- **Skill/Tool**: The user can select any of the options available in this field, as many as necessary but one at the time.
+- **Skill/Tool**: The user can select any of the options available in this field, as many as necessary but one at a time.
- **Description**: Read-only field. It shows the description of the tool, used by the agent to choose the appropriate tool for each case.
+- **Model**: This field appears only when the selected tool has the **Use Model** checkbox enabled in the [Skill/Tool window](#skilltool-window). It allows you to configure a specific LLM model for this tool in this agent. The model must be specified using the format `provider/modelname` (e.g., `openai/gpt-4`, `anthropic/claude-3-5-sonnet`). If left empty, a default model will be selected depending on the tool's implementation.
+- **Module**: Module in which this tool configuration will be exported. This field is only available with the `System Administrator` role.
- **Active**: checkbox to activate the tool.
!!!info
@@ -176,6 +181,14 @@ Fields to note:
In the Knowledge Base File window, you can define the files with which the agents can interact.
+!!!info "Image Indexing"
+ When files are indexed in an agent's knowledge base, **image files are handled differently** from text documents:
+
+ - **Text documents** (PDF, TXT, MD, etc.) are indexed in the main vector database for semantic search using the Knowledge Base Search tool
+ - **Image files** (PNG, JPG, JPEG, etc.) are indexed in a **separate image database** specifically designed for visual similarity search
+ - This image database is currently used by the [OCR Tool](../../developer-guide/etendo-copilot/available-tools/ocr-tool.md) to find reference templates with visual markers that guide data extraction
+ - Each agent maintains its own image database, separate from its text knowledge base
+
### Header

@@ -278,7 +291,7 @@ Fields to note:
Fields to note:
- - **OpenAPI Flow** Only show if the **OpenAPI Flow Specification** is chosen in the Type field. OpenAPI Flow selector, grouping enpoints common to a specific functionality.
+ - **OpenAPI Flow** Only show if the **OpenAPI Flow Specification** is chosen in the Type field. OpenAPI Flow selector, grouping endpoints common to a specific functionality.
=== "Remote File"
@@ -328,6 +341,10 @@ In this window , the user can find [available tools](../../developer-guide/etend

+Fields to note:
+
+- **Use Model**: Checkbox that indicates whether this tool requires an LLM model to function. When checked, a **Model** field will appear in the Skills and Tools tab of the Agent window, allowing you to configure a specific model for this tool in that agent.
+
Some tools require to communicate with Etendo through WebHooks. Their configuration can be found in the Webhooks tab.
!!!info
@@ -348,6 +365,25 @@ In this window, it is possible to configure access roles for each Agent. This me
!!!note
In case of deleting an agent, the related agent access records are also deleted.
+## Agent Memory Window
+
+:material-menu: `Application` > `Service` > `Copilot` > `Agent Memory`
+
+The Agent Memory window allows you to capture and reuse rules and knowledge acquired in any Copilot agent. Every memory you register is tied to a specific agent and is automatically injected into its answers according to your organization, role, and user context.
+
+
+
+Fields to note:
+
+- **Organization**: Defaults to the agent's organization; may be left blank for a global memory. Copilot only injects entries that belong to the current organization tree unless the value is empty.
+- **Active**: Enables or disables the memory without deleting it. Inactive rows never reach the conversation.
+- **User/Contact**: Optional user owner. Leave empty to expose it to everyone. Only the selected user sees the memory.
+- **Role**: Optional role filter. Any user working under the chosen role will receive the hint.
+- **Text Field**: The actual content Copilot will append. It is recommended to use short, action-oriented statements.
+
+!!! info
+ When multiple memories match, Copilot lists them as bullet points under “Use the following relevant previous information.”
+
## Process Request Window
:material-menu: `Application`>`General Setup`>`Process Scheduling`>`Process Request`
@@ -409,7 +445,7 @@ In this window, the user can find and add AI models to be used by the agents, Av
!!!info
- Automatically, the window will be populated with the Etendo default distributed models, after the first agent synchronization.
- - Also diffrent models and providers must be entered manually.
+ - Also different models and providers must be entered manually.

diff --git a/mkdocs.yml b/mkdocs.yml
index 4954f97f52..e760d09074 100644
--- a/mkdocs.yml
+++ b/mkdocs.yml
@@ -599,6 +599,7 @@ nav:
- How to Execute Copilot Through the Console: developer-guide/etendo-copilot/how-to-guides/how-to-execute-copilot-through-console.md
- How to Explore ERP Data with MCP: developer-guide/etendo-copilot/how-to-guides/how-to-explore-etendo-data-with-postgres-and-chart-mcp.md
- How to Export Tools and Agents: developer-guide/etendo-copilot/how-to-guides/how-to-export-tools-and-assistants.md
+ - How to Improve OCR Recognition: developer-guide/etendo-copilot/how-to-guides/how-to-improve-ocr-recognition.md
- How to Integrate Copilot with Jira via MCP: developer-guide/etendo-copilot/how-to-guides/how-to-integrate-copilot-with-jira-mcp.md
- How to Integrate a Sales Assistant with Shopify MCP: developer-guide/etendo-copilot/how-to-guides/how-to-integrate-sales-assistant-with-shopify-mcp.md
- How to Use an Agent as MCP Server: developer-guide/etendo-copilot/how-to-guides/how-to-use-an-agent-as-mcp-server.md