# Welcome to Decube

Reliable data, better decisions.

## AI Readiness starts with Trusted Data.

**Decube** is a comprehensive data trust platform designed for the AI era. We provide a unified solution for data observability, catalog, and governance that empowers organizations to build AI-ready data foundations with confidence.

## Observe. Discover. Govern. Trust - Unified approach

Our platform enables data teams to establish complete trust in their data through intelligent monitoring, seamless discovery, robust governance, and enterprise-grade security.

Here's what you can expect after using Decube:

1. **Foresee and Prevent**: Mitigate potential data issues before they escalate with ML-powered anomaly detection and real-time response capabilities.
2. **AI-Ready Data Foundation**: Get complete visibility and understanding of data going into AI/ML models to develop accurate models for real business impact.
3. **Accelerated Discovery**: Discover quickly the data you need to make informed decisions and break down data silos between teams with intelligent catalog and lineage.
4. **Enterprise-Grade Security**: Protect your sensitive and PII data by implementing strict access controls with SOC 2, ISO 27001, and HIPAA compliance.

{% embed url="<https://www.youtube.com/watch?v=IYby2oXSQlU>" %}

Follow our handy guide to get started as quickly as possible:

{% content-ref url="/pages/7arv6rYvHmkMYCoY8APK" %}
[Getting started](/overview/getting-started)
{% endcontent-ref %}

### Fundamentals: Dive a little deeper

<figure><img src="/files/vXBoOJnbyb7BxmIeg11l" alt=""><figcaption></figcaption></figure>

Learn the fundamentals of Decube to get a deeper understanding of our main features:

{% content-ref url="/pages/3SHtVpZ9FTA4QhLxNp6E" %}
[Incidents Overview](/data-quality/data-quality)
{% endcontent-ref %}

{% content-ref url="/pages/ceSNi4efmZaeCMwDsgwA" %}
[Overview of Asset Types](/catalog/overview-of-asset-types)
{% endcontent-ref %}

{% content-ref url="/pages/avKaKeCbdL7lyZgUZ0cN" %}
[Glossary, Category and Terms](/glossary/glossary-category-and-terms)
{% endcontent-ref %}

### Leaders in Data Observability & Data Catalog

Trusted by numerous data-driven companies worldwide, decube provides a mark of data integrity, crucial for compliance, decision-making, and operational insights in the AI era.

***

### Need Help?

If you have any questions that you can't find in this documentation, reach out to us:

* **Email**: [hello@decube.io](mailto:support@decube.io)
* **Request a Demo**: [Schedule a personalized demo](https://decube.io/request-a-demo)
* **Explore Sandbox**: [Try our interactive sandbox](https://www.decube.io/explore-sandbox)


# Getting started

Get your data trust platform up and running in 30-60 minutes.

## Before You Begin

**Time Required:** 30-60 minutes\
**Prerequisites:** Access to your data source (database, warehouse, etc.)

### What You'll Accomplish:

✅ Connect your first data source\
✅ View your data assets in the catalog\
✅ Set up basic data quality monitoring\
✅ Configure alert notifications

### What You'll Need:

* [ ] Credentials with the right permissions for your data source
* [ ] Network access permissions (if behind VPN/firewall)
* [ ] Email/Slack for notifications (optional)

{% hint style="info" %}
Make sure you're using the Decube login URL for your region (APAC, EU, or US).
{% endhint %}

### Role-Specific Guidance:

* **👨‍💻 Data Engineers:** Follow the technical setup steps and configuration guides
* **🛡️ Governance Teams:** Use this guide to understand what to request from your technical team and what outcomes to expect
* **👤 Admins:** Focus on user management and access control sections

## Quick Start: See Decube in Action

### 1. Login to Decube

Enter your credentials and click Login. You'll then be signed in and directed to Decube's landing page.

<figure><img src="/files/oKqSe5lsL7PMdjjxTcMD" alt=""><figcaption><p>Sign-in page</p></figcaption></figure>

{% hint style="info" %}
To reset your password, click Forgot Password on the login page and follow the instructions in the email you receive. If you don't see the email, please contact us at <hello@decube.io>.
{% endhint %}

<figure><img src="/files/2nNJdC7Hl03uBb4JtTRx" alt=""><figcaption><p>Forgot your password</p></figcaption></figure>

### 2. Connect Your Data Source (10-15 minutes)

**Why this matters:** This is the foundation - without a data source, Decube has nothing to monitor or catalog.

**👨‍💻 For Data Engineers:** Follow the technical setup guides below\
**🛡️ For Governance Teams:** Share these links with your technical team and ensure they have the necessary credentials

Connecting your first data source when you start with an account is compulsory. Each data source may have different prerequisites or configurations to be made; so we've prepared a guide on [**How to connect data sources**](/overview/getting-started/how-to-connect-data-sources) for each data source that you'd like to add.

<figure><img src="/files/KJyb5pNiBPgHXrNQs9Jt" alt=""><figcaption><p>Onboarding your data source with no code</p></figcaption></figure>

Connection to your data source should take a few minutes depending on the volume of your data. Once your data source has been connected, you can then navigate to the **Catalog** module.

✅ **Success indicator:** You see tables appearing in the Catalog module

### 3. Explore Your Data Catalog (5 minutes)

**Why this matters:** Verify your data is being ingested and cataloged correctly.

Navigate to **Catalog** → Browse your connected tables and schemas

**🛡️ For Governance Teams:** This is where you'll see all your data assets organized and searchable. You can add business context, classifications, and glossary terms here.

✅ **Success indicator:** You can see your database tables, schemas, and basic metadata

## Complete Setup: Configure Monitoring & Alerts

### 4. Enable Data Quality Monitoring (10-15 minutes)

**Why this matters:** Proactive monitoring prevents data issues from impacting downstream systems and AI/ML models.

Navigate to **Data Quality** → **Config** → Enable monitors for critical tables

**Quick Setup Options:**

* **Freshness Monitoring:** Detect when data stops updating
* **Volume Monitoring:** Catch unusual data volume changes
* **Schema Drift:** Alert on table structure changes

**👨‍💻 For Data Engineers:** Configure custom SQL monitors and set appropriate thresholds\
**🛡️ For Governance Teams:** Focus on business-critical tables and data lineage impact

[Detailed monitor setup guide →](/data-quality/enable-asset-monitoring)

✅ **Success indicator:** You see monitors active on your critical tables

### 5. Configure Alert Notifications (5-10 minutes)

**Why this matters:** Get notified immediately when data issues occur, enabling faster response times.

Navigate to **My Account** → **Config Settings** tab → Set up alert channels

**Alert Options:**

* **Email notifications:** Direct alerts to your inbox
* **Slack integration:** Team channel notifications
* **Microsoft Teams:** Collaborate on incident resolution
* **Webhook endpoints:** Custom integrations

**👤 For Admins:** Set up organization-wide notification policies\
**🛡️ For Governance Teams:** Ensure stakeholders are notified of data quality issues

[Alert setup guide →](/alert-notifications/notification-alerts)

<figure><img src="/files/AG2pTCFS5JMf4QZD0eEn" alt=""><figcaption><p>Add your alert channels in the Config Settings</p></figcaption></figure>

✅ **Success indicator:** You receive a test notification

### 6. Set Up User Access (10 minutes)

**Why this matters:** Ensure your team can access relevant data while maintaining security and compliance.

Navigate to **My Account** → **User Management**

**👤 For Admins:**

* Invite team members with appropriate roles
* Configure group-based access policies
* Set up data source permissions

**🛡️ For Governance Teams:**

* Define data classification policies
* Set up approval workflows for sensitive data access
* Configure glossary terms and business context

[User management guide →](/manage-access/user-management)\
[Group access controls →](/group-access-policies/groups-management-overview)

✅ **Success indicator:** Team members can log in and see appropriate data

## Congratulations! You're Set Up

### What Happens Next:

* **First 3 days:** Decube analyzes your data patterns and establishes baselines. ML models learn your data behavior for accurate anomaly detection.
* **Ongoing:** Receive proactive alerts for any data quality issues

### Explore Additional Features:

**Data Discovery & Catalog:**

{% content-ref url="/pages/ceSNi4efmZaeCMwDsgwA" %}
[Overview of Asset Types](/catalog/overview-of-asset-types)
{% endcontent-ref %}

**Business Glossary & Governance:**

{% content-ref url="/pages/avKaKeCbdL7lyZgUZ0cN" %}
[Glossary, Category and Terms](/glossary/glossary-category-and-terms)
{% endcontent-ref %}

**Bulk Metadata Management:** Use Export/Import to quickly update your Catalog metadata and Glossary terms in bulk through CSV workflows. Perfect for large-scale onboarding and governance updates.

{% content-ref url="/pages/y9ah9lgv3RE1udeSYB55" %}
[Export/Import Overview](/export-import/export-import-overview)
{% endcontent-ref %}

**Advanced Data Quality:**

{% content-ref url="/pages/3SHtVpZ9FTA4QhLxNp6E" %}
[Incidents Overview](/data-quality/data-quality)
{% endcontent-ref %}

**Public API (BETA):** For technical users: Integrate Decube with your existing tools and workflows using our REST API. Access data catalog, quality metrics, lineage, and user management programmatically.

{% content-ref url="/pages/UGzKctB56AdkVf37rgbd" %}
[Overview](/public-api/overview)
{% endcontent-ref %}

### Need Help?

If you encounter any issues during setup:

* **Email Support:** <hello@decube.io>
* **Documentation:** Search our knowledge base
* **Community:** Join our user community for best practices

***

**Next Steps:** Once your data source is connected and monitoring is active, explore our advanced features like the Governance module, Collaboration features and granular access controls to build a comprehensive data trust platform.


# How to connect data sources

Connect your data sources to start building your data trust foundation.

## Getting Started with Data Source Connections

**Time Required:** 5-15 minutes per data source\
**Prerequisites:** Admin credentials and network access to your data sources

{% embed url="<https://youtu.be/BlNvlLFDmro?si=1PDbBIxqBFijCWc>\_" %}

### What You'll Need Before Starting:

Review the specific connector documentation for your data source (see links below):

* [ ] Network connectivity (VPC access if needed)
* [ ] The right permissions assigned to the credentials you will use
* [ ] Connection details (host, port, database names etc. specific to your data source)

{% hint style="info" %}
If you have any questions or issues with connecting your data sources, please **initiate a live chat** with us from the bottom right of our app page and we'll get someone to walk you through the connection options.
{% endhint %}

## Recommended Connection Order

**🎯 Start Here:** We highly recommend connecting your **Data Warehouses** or **Relational Databases** first so you can see your tables immediately within Decube's Catalog module.

**Why this order?**

1. **Data Warehouses/Databases** → Immediate catalog visibility and data quality monitoring
2. **Transformation Tools** → Add lineage and pipeline monitoring
3. **Business Intelligence** → Complete end-to-end data observability
4. **Data Lake** → File-level governance

## Security & Network Access

If your data sources are not publicly accessible, you may need to allow Decube to access your VPC via secure methods:

{% content-ref url="/pages/QwJ2bt7KTcAr8WngRXFh" %}
[Enabling VPC Access](/security-and-connectivity/enabling-vpc-access)
{% endcontent-ref %}

{% content-ref url="/pages/IeabSxBfHBW69C0lrCKD" %}
[IP Whitelisting](/security-and-connectivity/ip-whitelisting)
{% endcontent-ref %}

{% content-ref url="/pages/jU79kSLrPa7VgOtHKw6b" %}
[SSH Tunneling](/security-and-connectivity/ssh-tunneling)
{% endcontent-ref %}

Below are quick links to each data source we support.

### Data Warehouses

* [Snowflake](/warehouses/snowflake)
* [Redshift](/warehouses/redshift)
* [Google Big Query](/warehouses/google-big-query)
* [Databricks](/warehouses/databricks)
* [Azure Synapse](/warehouses/azure-synapse)
* [ClickHouse](/warehouses/clickhouse)
* [Microsoft Fabric](/warehouses/microsoft-fabric)

### Relational Databases

* [PostgreSQL](/databases/postgresql)
* [MySQL](/databases/mysql)
* [SingleStore](/databases/singlestore)
* [Microsoft SQL Server](/databases/microsoft-sql-server)
* [Oracle](/databases/oracle)
* [SAP HANA](/databases/sap-hana)

### Transformation tools

* [dbt](/transformation-tools/dbt)
* [dbt Core](/transformation-tools/dbt-core)
* [Fivetran](/transformation-tools/fivetran)
* [Airflow](/transformation-tools/airflow)
* [AWS Glue](/transformation-tools/aws-glue)
* [Azure Data Factory](/transformation-tools/azure-data-factory)
* [Apache Spark](/transformation-tools/apache-spark)
* [OpenLineage](/transformation-tools/openlineage)

### Business Intelligence

* [Tableau](/business-intelligence/tableau)
* [Looker](/business-intelligence/looker)
* [PowerBI](/business-intelligence/powerbi)

### Data Lake

* [AWS S3](/datalake/s3)
* [Azure Data Lake Storage (ADLS)](/datalake/azure-data-lake-storage-adls)
* [Google Cloud Storage (GCS)](/datalake/google-cloud-storage-gcs)

### SaaS

* [Salesforce](/saas/salesforce)


# Changelog

Here's what's new, improved, and fixed in each Decube release.

{% hint style="info" %}
Want these delivered to your inbox instead of checking back here? **Subscribe in the link below** for email updates ↓
{% endhint %}

{% embed url="<https://preview.mailerlite.io/forms/2072223/192390591821120552/share>" %}

***

{% updates format="full" %}
{% update date="2026-08-13" %}

## 1.60.12: Fixes and Improvements

#### Improvements

* **\[Catalog]** Added a DATETIME\_TZ unified column type so time-zone-aware columns are collected and displayed correctly.
* **\[Lineage]** SQL lineage now checks that both assets exist before an edge is written.
* **\[Lineage]** Lineage now retries automatically when a network request fails to load the graph, and shows an error in the UI if the retry does not succeed.

#### Bug Fixes

* **\[Catalog]** Fixed an issue where the catalog listing page crashed when a filtered custom attribute was deleted.
* **\[Catalog]** Fixed an issue where Data Product mention links used the wrong URL and returned a 500 error.
* **\[User Management]** Fixed an issue where sorting by status did not group users correctly.
* **\[Data Source Management]** Fixed an issue where the **PowerBI** extractor silently dropped workspace batches and report lineage.
* **\[Data Source Management]** Fixed an issue where Airflow metadata collection intermittently failed with a 503 "no healthy upstream" error.
  {% endupdate %}

{% update date="2026-08-10" %}

## 1.60.11: Fixes and Improvements

#### Bug Fixes

* **\[Data Source Management]** Fixed an issue where serverless **Synapse** sources could not be profiled.
* **\[Asset Details]** Fixed an issue where the yellow highlight was missing when a change request only modified linked terms.
  {% endupdate %}

{% update date="2026-08-07" %}

## 1.60.10: Incident Remediation with Trusty

#### New Features & Enhancements

* **\[AI Hub]** Trusty AI can now suggest remediation steps on incidents.
* **\[AI Hub]** Refreshed the AI Hub interface.

#### Bug Fixes

* **\[Asset Details]** Fixed an issue where filters in the Incidents and Monitors tabs persisted when you switched between assets.
* **\[Asset Details]** Fixed an issue where the Feed tab displayed the wrong icon.
* **\[UI]** Fixed an issue where calendar dates rendered one day earlier for viewers in negative-UTC-offset timezones.
* **\[Dashboard]** Fixed an issue where source-level requests were missing from the Recent panels on limited access accounts.
  {% endupdate %}

{% update date="2026-08-03" %}

## 1.60.9: AI Agents Registry (Early Access) & Term Linking

#### New Features & Enhancements

* **\[AI Hub]** Consolidated Trusty and the new AI Agents Registry into a single AI Hub on the navbar.
* **\[AI Hub]** Introduced the [AI Agents Registry](/ai-agent-registry/overview) (Limited preview). This allows you to register agents in your Decube organization, with automated registration enabled for Snowflake cortex. Manual registration is also supported.
* **\[Asset Details]** You can now [link glossary terms directly from the Asset Details page](/catalog/data-catalog/updating-asset-details#linking-glossary-terms).

#### Bug Fixes

* **\[Governance]** Fixed an issue where the data source icon was missing in the Recent Requests panel.
* **\[Lineage]** Fixed an issue where duplicated data jobs appeared in the Connection Details sidebar in Manual Lineage.
* **\[Tableau]** Fixed an issue where **Tableau** executions were canceled due to timeout.
* **\[Monitors]** Fixed an issue where the Back and Submit buttons were missing when creating a Custom SQL Monitor.
  {% endupdate %}

{% update date="2026-07-30" %}

## 1.60.8: Trusty Skills & Lineage Visuals

#### New Features & Enhancements

* **\[Trusty]** Added Skills support to Trusty, letting it run guided, multi-step workflows on your data assets.
* **\[Lineage]** Lineage graph nodes now use type-based visuals to distinguish asset types at a glance.

<figure><img src="/files/XV75tzrgFZ4Bk64yZZKv" alt=""><figcaption></figcaption></figure>

#### Improvements

* **\[Glossary - Linked Assets]** Linking assets from the Glossary now has a new look, making it easier to see all selected assets and curate your linked assets in one view.
* **\[Asset Cards]** Streamlined all places where asset cards, or asset names with their source names, are shown.

#### Bug Fixes

* **\[User Management]** Fixed an issue where Entra sign-in failures caused an infinite error loop instead of redirecting to the login page.
* **\[Lineage]** Fixed an issue where the lineage endpoint ran out of memory on datasets with very large graphs.
* **\[Lineage]** Fixed an issue where long source names pushed the asset type dropdown out of view in the Manual Lineage Modal.
* **\[Catalog]** Fixed extra spacing before the condition in the attributes filter.
* **\[Glossary]** Fixed an issue where newly created terms did not inherit the parent glossary in change requests.
  {% endupdate %}

{% update date="2026-07-27" %}

## 1.60.7: Virtual Sources Icon Library & Search Improvements

#### Improvements

* **\[Search]** Search experience is improved to more likely return exact-matching asset names as the first few results on the search result.
* **\[Data Source Management]** You can now select from a library of icons to replace how the virtual source is represented in your catalog.
* **\[Data Source Management]** Updated the **SAP HANA** connector logo icon.

#### Bug Fixes

* **\[Data Recon]** Fixed an issue where source names weren't properly truncated on the Target Table view.
* **\[My Account]** Fixed an issue where asset names were cut off in the Source-based Policy modal under Group Management.
* **\[Catalog]** Fixed an issue where deleted-source assets in the Asset Picker incorrectly displayed an "Access restricted to this asset" message.
* **\[Data Source Management]** Fixed an issue where scheduler scan periods were incorrect on servers in non-UTC timezones.
* **\[Data Governance]** Fixed an issue where the DQ scorecard report failed when row counts were null.
* **\[Data Source Management]** Fixed an issue where **SAP HANA** connections could time out.

{% embed url="<https://youtu.be/mHf-6NWuRnE?si=KvYr75dKhlhmucOR>" %}
{% endupdate %}

{% update date="2026-07-20" %}

## 1.60.6: New SAP HANA Connector

#### New Features & Enhancements

* **\[Data Source Management]** **SAP HANA** connector (beta): Connect your SAP HANA database to Decube to automatically catalog its metadata.

#### Improvements

* **\[Data Source Management]** Improved the **Looker** connector.
* **\[Data Source Management]** Improved the **Fivetran** connector.
* **\[Trusty]** Approval workflows now support queuing multiple approval requests.

#### Bug Fixes

* **\[Incidents]** Fixed an issue where fetching the Jira issue type returned an error.
* **\[Data Source Management]** Fixed an issue where **dbt Cloud** metadata jobs returned 404 errors for URLs containing a double slash.
* **\[Search]** Increased the reindex timeout to 45 minutes to prevent failures on large reindex jobs.
  {% endupdate %}

{% update date="2026-07-16" %}

## 1.60.5: Trusty Visualisations & Monitor Creation

#### New Features & Enhancements

* **\[Trusty]** Visualisations: Trusty detects when an answer is best expressed visually and renders it inline in the conversation. Use cases such as lineage, pipelines, and relationships can be rendered as diagrams that you can copy and download.
* **\[Trusty]** Monitor Creation: Trusty proposes a monitor with a recommended test type and configuration, and creates it only after you approve.

#### Bug Fixes

* **\[Glossary]** Fixed an issue where updating Related Terms could trigger an infinite render loop.
* **\[Asset Details]** Fixed an issue where linked Data Products incorrectly showed a "Get access" prompt.
* **\[Monitors]** Fixed an issue where Freshness monitor history returned 500 errors on on-demand tests.
* **\[Export/Import]** Fixed an issue where importing catalog assets with column-level Custom Attributes failed.
* **\[Lineage]** Fixed an issue where **Tableau** lineage introduced incorrect column identifiers after PublisherV3.
* **\[Data Source Management]** Fixed an issue where **Databricks** v3 task runs proceeded on stale data instead of being skipped.
  {% endupdate %}

{% update date="2026-07-14" %}

## 1.60.4: Decube MCP Fixes

#### Bug Fixes

* **\[MCP]** Fixed an issue where the `create_monitor` MCP tool rejected requests with `test_type` set to null.
* **\[MCP]** Fixed an issue where the `update_custom_attribute` MCP tool rejected integer values (e.g. 2, 3).
* **\[MCP]** Fixed an issue where the `get_profiling_result` MCP tool returned raw ORM/Pydantic objects instead of serialized data.
  {% endupdate %}

{% update date="2026-07-13" %}

## 1.60.3: Trusty QoL & UX Improvements

#### Feature Release: Trusty QoL & UX Improvements

Trusty just got several quality-of-life and UX upgrades to make conversations faster and easier to work with:

* **Follow-up prompts**: Depending on the conversation, Trusty may suggest follow-up prompts you can click to send instantly as a reply.
* **Copy**: Copy code snippets (SQL, .py, .json) and tables inline, in formats suited to the artifact (e.g. **.csv** for tables), or copy an entire response as text.
* **Download**: Download artifacts in formats suited to their content, such as **.csv** or **Markdown** for tables.
* **Maximise**: Expand certain artifacts into a larger modal view.
* **Regenerate**: Regenerate a response, appended to the bottom of the thread.
* **Scroll to latest**: A "scroll to latest" button appears when you're midway through a long conversation with content below the fold.
* Tool calls are now shown while a response is being generated.

#### Bug Fixes

* **\[AI]** Fixed an issue where Cursor could not integrate with the MCP Server.
* **\[AI]** Fixed an issue where error logs leaked into MCP responses.
* **\[Data Source Management]** Fixed an issue where the **AzureSQL** connection form was missing a port field.
  {% endupdate %}

{% update date="2026-07-09" %}

## 1.60.2: Decube MCP Actions

### Improvements

* \[AI] Added write and action tools to the MCP Server, letting connected AI tools take actions in Decube.
* \[Data Source Management] Added placeholder text to source form inputs.

### Bug Fixes

* \[Catalog] Fixed an issue where the Feed scrollbar flickered.
* \[Data Reconciliation] Fixed an issue where a stale label persisted on the Configuration tab.
* \[My Account] Fixed an issue where long virtual source names were not truncated.
* \[Access Requests] Fixed an issue where implicit ACL policies were not auto-submitted when access was granted.
* \[Data Source Management] Fixed an issue where the data source card title shrank the owners display.
* \[Modals] Fixed an issue where the Test Type filter dropdown flickered when opening and closing.
* \[Export/Import] Fixed an issue where importing catalog assets with linked terms failed.
  {% endupdate %}

{% update date="2026-07-06" %}

## 1.60.1: MCP Server

### Feature Release: MCP Server

[MCP Server is now available](/mcp/overview), letting you connect external AI tools and agents to Decube through the Model Context Protocol using OAuth authentication.

### Improvements

* \[Data Source Management] Source connection forms now validate on submission instead of as you type.

### Bug Fixes

* \[Data Source Management] Fixed issues affecting Fabric metadata collection following its recent release.
  {% endupdate %}

{% update date="2026-07-02" %}

## 1.60.0: Data Products

### Feature Release: Data Products

[Data Products is now available](/data-products/overview). It's a curation layer on top of your catalog that lets producers package trusted, documented data products and lets consumers discover, evaluate, and use them with confidence — all without leaving Decube. See the Data Products overview for details.

* ​\[Data Products/Creation] Use the creation wizard to package any catalog asset into a named, documented data product with an assigned owner. Owners can manage assets, update descriptions, and track usage.
* ​\[Data Products/Marketplace] Browse the marketplace to search, filter, and sort published data products. Each product card shows its description, asset count, peer star rating, and consumer count, with a detail page for documentation, questions, and ratings.
* ​\[Data Products/Governance] Admins can control who can create, edit, or delete data products through module-based access control. Ownership is always tracked, and products are visible to your whole organization by default.

### New Features & Enhancements

* ​\[Data Source Management] Added support for Fabric as a data source. See the Microsoft Fabric setup guide.

### Improvements

* \[Lineage] Added expiration handling for SQL query-based lineage records.
* \[Data Source Management] Improved Snowflake SQL query history lookup performance.
* \[UI] Updated status indicator sizing across components.

### Bug Fixes

* \[Search] Fixed an issue where the catalog search input lagged.
* ​\[Catalog/Asset Card] Fixed an issue where overflowing tooltips were cut off in asset cards.
* ​\[My Account/Group Management] Fixed an issue where descriptions were cut off when user lists wrapped in Group Management.
* ​\[Catalog/Requests] Fixed an issue where the footer cut off the bottom of scrollable content in Catalog Requests.
* ​\[Lineage] Fixed an issue where source and target columns displayed "\_" for virtual columns in Connection Details.
* ​\[My Account/Group Management] Fixed content misalignment in the Business Glossary ACL modal.
* ​\[User Management] Fixed an issue where rapid Sign In button clicks caused an authentication loop.
* ​\[Data Source Management] Fixed UI text overflow on the Sources page.
* ​\[Data Source Management] Fixed an issue where dbt Cloud metadata collection expected timestamps to always include microseconds.
  {% endupdate %}
  {% endupdates %}

## \[Release Version: 1.59.2] 29 Jun 2026

**\[Improvements]**

* **\[Data Source Management]** Improved **BigQuery** connector support.
* **\[Data Recon]** Optimized recon run scheduling to distribute long-running recon runs more evenly across tasks, reducing overall execution time.

## \[Release Version: 1.59.1] 25 Jun 2026

**\[Bug Fixes]**

* **\[Trusty]** Fixed an issue where the Trusty AI service hit the Elasticsearch scroll context limit, causing 429 errors.

## \[Release Version: 1.59.0] 22 Jun 2026

**\[New Features & Enhancements]**

* **\[Data Source Management]** Virtual Sources and Objects are now available, letting you define and manage virtual data sources and objects. Check out the documentation for more details: [Virtual Sources and Objects](/virtual-sources-and-objects/virtual-sources).

**\[Bug Fixes]**

* **\[User Management]** Fixed incorrect copy in the confirmation modal.
* **\[Catalog]** Fixed an issue where applied filters were retained when switching between accounts.
* **\[Incidents]** Fixed an issue where the page scrollbar toggled when interacting with the test type filter dropdown.
* **\[Lineage]** Fixed an issue where expanding a node removed the expand button on nearby nodes.
* **\[Data Recon]** Fixed an issue where Data Recon jobs remained stuck in a queued status.

## \[Release Version: 1.58.1] 18 Jun 2026

**\[Bug Fixes]**

* **\[Lineage]** Fixed an issue where asset names in search results were not truncated in the Incident Details and Lineage views.
* **\[Data Source Management]** Fixed an issue with **Redshift** column type change handling during metadata collection.

## \[Release Version: 1.58.0] 15 Jun 2026

**\[Improvements]**

* **\[Asset Details]** The legacy Tags field has been removed from assets. Tags are now managed through Custom Attributes.
* **\[Data Source Management]** Updated the **PostgreSQL** metadata collector to the V3 collection architecture.
* **\[Data Source Management]** Updated the **Redshift** metadata collector to the V3 collection architecture.

**\[Bug Fixes]**

* **\[Incidents]** Fixed an issue where the default incident listing excluded incidents created after the API process started. The default time window is now computed per request.
* **\[Data Source Management]** Fixed **dbt Cloud** metadata collection failing with a 502 error on large job payloads.
* **\[Data Source Management]** Fixed an issue where **ADLS** single-file discovery returned 0 tables for root-level wildcard path specs.
* **\[My Account/Group Management]** Fixed inconsistent checkbox alignment in the Data Recon section of source-based group policies.

## \[Release Version: 1.57.2] 11 Jun 2026

**\[Public API Update]**

* **\[Public API]** Added a new get Lineage API to the Public API.

**\[Improvements]**

* **\[Authentication]** Signing in now redirects you back to the original URL you were trying to visit.
* **\[Profiler]** Profiler now displays status indicators for profiling runs.

**\[Bug Fixes]**

* **\[Glossary]** Fixed a naming inconsistency in the Edit Linked Assets panel.
* **\[Monitors]** Fixed a 422 error when switching a monitor's test type to Freshness.
* **\[Data Source Management]** Fixed a crash in the **S3** V3 metadata collector when a file was removed mid-collection.
* **\[Profiler]** Fixed a **Synapse** profiler crash (ValidationError) when the unique percentage exceeded 1.0.

## \[Release Version: 1.57.1] 4 Jun 2026

**\[New Features & Enhancements]**

* **\[Profiler]** You can now delete profiler results directly from the Profiler.
* **\[My Account/Group Management]** You can now grant non-Owner users permission to export and import metadata via the new **Allow Export/Import** administrative policy.
* **\[Org Settings]** Microsoft Entra SSO now supports authentication scoped to a specific tenant.

**\[Improvements]**

* **\[My Account/Group Management]** Simplified asset selection and permission model on source-based policies.
* **\[Data Source Management]** Additional config for data sources now includes a deselect option for sources.

**\[Bug Fixes]**

* **\[Data Source Management]** Fixed an issue with **PowerBI** config endpoint performance and connection clearing.
* **\[Monitors]** Fixed an issue where deleting a monitor returned a 500 error due to an Elasticsearch version conflict when the document was missing.
* **\[Profiler]** Fixed an **Oracle** profiler crash (ORA-01839) caused by a month-end relative date filter.
* **\[Data Source Management]** Fixed a **dbt Core** job yielding issue that could cause monitor data to be lost.

## \[Release Version: 1.57.0] 28 May 2026

**\[Improvements]**

* **\[Search]** Refactored search across frontend and backend for improved reliability and performance.
* **\[Modals]** Updated modal components to align with the design system.

**\[Bug Fixes]**

* **\[Lineage]** Fixed an issue where view lineage extraction failed with a KeyError on table alias lookup.
* **\[Incidents]** Fixed an issue where referencing a deleted column in Incident Debugger caused an unhandled error.

## \[Release Version: 1.56.7] 25 May 2026

**\[Improvements]**

* **\[Lineage]** The SQL statement textarea in the lineage editor now uses a text cursor, matching expected editing behaviour.
* **\[Monitors]** Removed the character limit on monitor descriptions.

**\[Bug Fixes]**

* **\[Search]** Fixed an issue where the public search API returned an error when results exceeded the maximum result window.
* **\[Lineage]** Fixed an issue where attribution lineage was not deleted when a data warehouse source was removed.
* **\[Governance]** Fixed an issue where change requests were sorting inconsistently.
* **\[Incident Details]** Fixed an issue where the sensitivity endpoint was called before the authentication check completed.
* **\[Data Source Management]** Fixed column deletion cascade and data retention for **S3** data lake sources.

## \[Release Version: 1.56.6] 21 May 2026

**\[Improvements]**

* **\[Catalog]** Pressing Enter in the Dashboard search bar now routes you to Catalog search.
* **\[Lineage]** The lineage canvas now lets you zoom out further to view large lineages in a single screen. Pan-on-scroll support means you can navigate using two-finger scroll on trackpads.
* **\[Profiler]** Input fields in the Profiler now display inline error states when validation fails.
* **\[Export/Import]** Optimised export query performance for faster data exports.

**\[Bug Fixes]**

* **\[Catalog/Terms]** Fixed an issue where terms were displaying incorrect Asset Documentation in the Entercard.
* **\[Data Source Management]** Fixed an issue with **Redshift** Late Binding Views causing ingestion errors.
* **\[Data Source Management]** Fixed a **PowerBI** extraction issue affecting unreferenced assets.

## \[Release Version: 1.56.5] 18 May 2026

**\[Improvements]**

* **\[UI]** Updated the sign-in message for improved clarity and user trust.
* **\[Incidents]** Updated copy text in the Jira integration panel within Incident Details.

**\[Bug Fixes]**

* **\[UI]** Fixed an issue where the horizontal tab indicator lost its blue colour when zooming in or out in the browser.

## \[Release Version: 1.56.4] 14 May 2026

**\[New Features & Enhancements]**

* **\[Catalog]** **Snowflake** constraints — including primary key, not null, unique, and foreign key — are now visible on each column in the Schema tab.
* **\[Reports]** Redesigned the Generate Report form, making required and optional filters and fields clearly distinct.
* **\[Filters]** Date filters are now datetime filters, allowing more precise filtering on date and time values.

**\[Improvements]**

* **\[Data Source Management]** "Test Connection" on Azure sources now returns a specific error message when the server connection is blocked by a firewall.
* **\[UI]** Updated floating toast notifications for UI consistency.

**\[Bug Fixes]**

* **\[Catalog/Terms]** Fixed a broken linked assets card on the Catalog request page when viewing linked terms.
* **\[Incidents]** Fixed a case count mismatch between the Lineage and Incident tabs.

## \[Release Version: 1.56.3] 12 May 2026

**\[New Features & Enhancements]**

* **\[Glossary]** Glossary terms now display a "Created at" timestamp.
* **\[Schema Tab]** Policy badges on the Schema tab now distinguish between policies applied directly and those applied through a classification rule.
* **\[Glossary]** The Overview and Custom Attributes tabs are now consolidated into a single view.
* **\[UI]** Standardised the date/time range picker component across multiple places in the app.

**\[Improvements]**

* **\[UI]** Standardised inline toast notifications across the app using a shared design system component.

**\[Bug Fixes]**

* **\[Data Source Management]** Fixed an issue with the **OpenLineage** and **Spark** source connector form.
* **\[Incidents]** Fixed an issue where charts were misaligned at the default Mac screen resolution.
* **\[Incidents]** Fixed an issue where the Jira user selector failed when the assigned user was null.
* **\[Data Source Management]** Fixed an issue where **PowerBI** sibling matching caused a deserialization error.
* **\[Incidents]** Fixed an issue where the incidents tab scrollbar flickered on Linux and Windows.

## \[Release Version: 1.56.2] 7 May 2026

**\[New Features & Enhancements]**

* **\[Incidents]** You can now view and edit monitors directly from the Incident Details page.
* **\[Lineage]** You can now download lineage results as a CSV directly from the Lineage view.
* **\[Catalog]** Added a Trusty Profiling Tool, enabling Trusty to surface profiling insights within the AI assistant.

**\[Improvements]**

* **\[Lineage]** Backwards-directing lineage edges now render under nodes for a cleaner graph layout.

**\[Bug Fixes]**

* **\[Incidents]** Fixed an issue where the Create Jira Issue button overlapped the Lineage Connection Details panel.
* **\[Incidents]** Fixed a Firefox issue where table row heights were not applied consistently, causing cell values to overlap.
* **\[Search]** Fixed an issue where non-persistent search queries were not reset when you exited a flow.

## \[Release Version: 1.56.1] 4 May 2026

**\[New Features & Enhancements]**

* **\[Incidents]** The Incident Details page now includes a Lineage component, letting you view data lineage directly from an incident.

**\[Improvements]**

* **\[Lineage]** In the column selection workflow, the "Clear All" button now only appears on the active node. Dropdowns on other nodes are removed to reduce confusion.

**\[Bug Fixes]**

* **\[Data Source Management]** Fixed an issue where **PowerBI** Native SQL failed to retrieve source configuration correctly.

## \[Release Version: 1.56.0] 29 April 2026

**\[Feature Release: Profiler Configuration]**

You can now configure what gets profiled before a run starts — scope to specific columns, filter rows by a date range, and choose how the data is sampled.

* Added three run options to **Generate a new profile**: **Quick Run** (system defaults), **Run with latest settings** (repeats the most recent configuration for that asset), and **Configure profile run** (opens a configuration modal).
* **Column selection** — choose to profile all columns or a specific subset using a searchable dropdown. Useful for wide tables or when you want to reduce profiling cost by scoping to relevant columns only.
* **Date/time filter** — restrict profiled rows to a fixed date range or a rolling relative window (e.g. Last 7 Days). Only available on assets with date or datetime columns.
* **Sampling strategy** — choose from Auto, Sampling Percentage, Target Row Count, or Full when no date filter is applied. When a date filter is active, options switch to All matching rows or Limit row count. Available options vary by source — see the source support matrix in the docs.
* Run settings are now shown on the profile results panel and on each entry in the historical profiles list.
* Added a **Download CSV** button to the profile results page to export profiling results.

Read more in the documentation: [Configure a profile run](/catalog/profiler/configure-profile-run).

## \[Release Version: 1.55.5] 27 Apr 2026

**\[Bug Fixes]**

* **\[Schema Tab]** Fixed an issue where a double scrollbar appeared when viewing column-level custom attributes.

## \[Release Version: 1.55.4] 23 Apr 2026

**\[New Features & Enhancements]**

* **\[Lineage]** Lineage columns now display with a link visual, making navigable connections easier to identify.

**\[Bug Fixes]**

* **\[Glossary]** Fixed an issue where the middle panel did not expand to fill available space.
* **\[Catalog/Asset Card]** Fixed an issue where classification tooltips overlapped the asset card.
* **\[Incidents]** Fixed an issue where the schema drift chart did not load.
* **\[Metrics]** Fixed an issue where **Databricks** schema names with quoted identifiers were not parsed correctly.

## \[Release Version: 1.55.2] 20 Apr 2026

**\[Feature Release: Observability & Governance Layers on the Lineage Graph]**

You can now overlay data quality and governance context directly on the lineage canvas — without leaving your current view. A new floating toolbar at the bottom of the **Lineage** tab gives you access to two independent layers that highlight the information most relevant to your role.

**Observability layer — surface incidents in context**

* Toggle **Show Incidents (Data Quality)** to highlight every table on the canvas that has open or muted incidents from the last 30 days.

**Governance layer — trace classification coverage across your pipeline**

* Toggle **Show Governance (Classifications)** to see aggregated classification counts on every node that contains classified columns. Hover over any badge to see a breakdown of which classifications are present.

Both layers can be active simultaneously for a combined view of data health and classification sensitivity on the same canvas.

Read more in the documentation: [Observability & Governance Layers](/lineage/automated-lineage/observability-and-governance-layers).

**\[New Features & Enhancements]**

* **\[Data Source Management]** Source configuration forms now include field validation, helping you catch errors before saving.

**\[Improvements]**

* **\[User Info]** The asset tree now loads only when you expand a collapsible section, improving initial page load time.

**\[Bug Fixes]**

* **\[Health Score]** Fixed a copy mistake in the Health Score section.
* **\[Export/Import]** Fixed incorrect capitalisation of "template" in the Import/Export module.
* **\[Profiler - MS SQL]** Fixed an integer overflow issue in MSSQL sources that caused incorrect handling of large row count values.

## \[Release Version: 1.55.0] 13 Apr 2026

**\[New Features & Enhancements]**

* **\[Lineage]** The Data Source field in the Add Connection form now supports search, making it easier to find a specific source.
* **\[Catalog]** Table headers now stay frozen while scrolling, keeping column labels visible as you navigate large datasets.
* **\[Catalog]** "Collection" now shows as "Schema", "Datasets" now show as "Table" and "Property" now shows as "Column" in the filter.
* **\[Asset Details]** Virtual columns now display their own distinct icon and tooltip in the asset tag view.

**\[Bug Fixes]**

* **\[Asset Details/Overview]** Fixed an issue where the "There is no more data to load" message was not centered on the Overview page.

## \[Release Version: 1.54.7] 9 Apr 2026

**\[Improvements]**

* **\[Data Source Management]** Clicking "Test Connection" on **ADF**, **Fivetran**, **Glue**, and **ADLS** sources now returns specific error messages for connection failures.
* **\[Lineage]** Improved lineage loading performance in Asset Details.
* **\[Governance]** Enhanced visual change tracking to show previous and current changes in the change request review workflow.
* **\[Modals]** Clicking outside configuration modals no longer accidentally dismisses them.
* **\[My Account]** Updated the icon for the Import tab in the Import/Export section.
* **\[Documentation]** Updated heading text copy across the app.

**\[Bug Fixes]**

* **\[My Account/User Info]** Fixed an issue where the Source Policy Permissions modal did not display permissions on load.

## \[Release Version: 1.54.5] 6 Apr 2026

**\[Improvements]**

* **\[Data Source Management]** Clicking "Test Connection" on **PowerBI**, **Synapse**, **dbt Core**, and **dbt Cloud** sources now returns specific error messages for connection failures.

**\[Bug Fixes]**

* **\[Incidents]** Fixed an incorrect logic check in the incident debug statement that could cause unexpected behaviour.

## \[Release Version: 1.54.4] 2 Apr 2026

**\[Improvements]**

* **\[Data Source Management]** When adding or modifying sources, clicking "Test Connection" on BigQuery, S3, Postgres, SQL Server, Oracle, Databricks, MySQL, Tableau, and Redshift now returns specific error messages for credential issues.
* **\[Lineage]** The search bar in Lineage now displays "Search to focus on node" to clearly explain the expected behaviour.

**\[Bug Fixes]**

* **\[Catalog/Requests]** When you close the modal without clicking "Save Preferences", filter values are no longer saved.
* **\[Search]** Fixed a behaviour where spaces and special characters were being stripped from search terms sent to server-side searches.
* **\[Catalog/Terms]** Fixed an issue where the Related Term component was still displaying the change request value after a change request was submitted.

## \[Release Version: 1.54.3] 30 Mar 2026

**\[Improvements]**

* **\[Documentation Component]** Added support for functional hyperlinks within the Documentation sections of both Asset Details and the Glossary.

**\[Bug Fixes]**

* **\[PowerBI Integration]** Resolved an issue where the PowerBI connection form would fail to save if no workspaces had been created yet. The form now saves successfully regardless of existing workspace count.

## \[Release Version: 1.54.2] 26 Mar 2026

**\[New Features & Enhancements]**

* **\[Lineage]** Added search functionality to the lineage canvas.
* **\[Lineage]** Added exclusion filters to the OpenLineage connector to exclude certain file paths from lineage generation.
* **\[Schema Tab]** Implemented improved search logic for the Schema tab in Asset Details.
* **\[Governance]** Updated UI copy for optional fields to improve clarity.

**\[Bug Fixes]**

* **\[User Management]** Fixed an issue where asset names displayed extra whitespace in the granted permission modal.
* **\[Catalog/Requests]** Resolved a navigation bug where you were not redirected after approving or rejecting a request from the summary modal.
* **\[Asset Details]** Fixed a functional regression where the "Request Access" modal failed to trigger correctly.
* **\[Catalog/Requests]** Fixed a CSS clipping issue where the Reviewer status dropdown was hidden by the table container.
* **\[Catalog/Asset Card]** Corrected missing padding and spacing next to the Business Owner icon.
* **\[Glossary]** Fixed a UI bug where the "Edit Linked Asset" button was clipped or hidden during window resizing.
* **\[Glossary/Term]** Fixed a logic error where "Request Changes" was enabled even when no dropdown changes had been made.

## \[Release Version: 1.53.0] 17 Mar 2026

**\[Feature Release: New Profiler]**

The new profiler is here! With a redesigned UI and enhanced capabilities, it provides deeper insights into your data quality and health. Key features include:

* Historical profilers: Access past profiling runs to track changes in your data over time.
* Table and column-level statistics: Get detailed metrics on your datasets, including row counts, null percentages, unique counts, and more.

Read more about the new capabilities and source support in the documentation: [Profiler](/catalog/profiler) and [Profiler source support and limitations](/catalog/profiler/profiler-source-support-and-limitations).

**\[Enhancement: Visual change tracking]**

* When changes are made in the Asset Details or Glossary, when they are being reviewed by the selected reviewer, they can see the tabs that have changes by the yellow dot highlights.
* Reviewers can see the before/after changes on the documentation by toggling between the states.

<figure><img src="/files/WfH7GicsfdkZPBYI6nr6" alt=""><figcaption><p>Highlighting of tabs that have changes.</p></figcaption></figure>

<figure><img src="/files/5wsGXIs4q55UvN2qc4Ty" alt=""><figcaption><p>Documentation showing current version, which you can toggle to see the previous.</p></figcaption></figure>

<figure><img src="/files/xoL8ojxQqvduoWYPOenB" alt=""><figcaption><p>Documentation showing previous version.</p></figcaption></figure>

## \[Release Version: 1.52.0] 12 March 2026

**\[Feature Release: Trusty AI]**

**Trusty AI is available for BETA access!**

1. **Conversational Metadata Assistant:** Trusty AI allows you to query your data catalog using natural language. Instead of manually browsing tables, you can now ask questions to surface assets, check data health, and perform impact analysis instantly.
2. **Semantic Search:** Previously, users had to search for exact technical table names. This has been updated so that Trusty understands business intent, allowing you to find relevant datasets (e.g., "Sales") even if the technical names are different.
3. **Data Health Summary:** Determining if a table is "safe" previously required checking multiple tabs for freshness and null rates. Trusty now provides a "Fitness for Use" summary in a single response, synthesizing monitor logs and profiling metrics.
4. **Impact Analyzer:** To prevent breaking changes, engineers previously traced lineage graphs manually. Trusty now automates this via chat, identifying downstream dependencies and potential risks before you deploy code changes.

If you are interested to participate in the BETA program, please reach out to your dedicated Account Manager for us to assess your account's eligibility for this phase.

Read more on [Trusty AI: Overview](/trusty-ai/trusty-ai) and [Suggested Questions to ask Trusty](/trusty-ai/trusty-ai/suggested-questions-to-ask-trusty).

**\[Feature Release: Bulk Update Incident Status]**

**What's new:** select up to 1,000 incidents and bulk update them to a single target status.

<figure><img src="/files/oJ1SmsBVwwOlxlILPQFA" alt=""><figcaption></figcaption></figure>

Managing incidents one by one, specifically updating their statuses can be tedious especially when noise spikes and many of the open alerts need the same action. You can now select and update multiple incidents at once directly from the Incident Overview page.

<figure><img src="/files/CVdidmF6pgAGI1KiBggG" alt=""><figcaption></figcaption></figure>

**Mixed Status Selection Handling:**

* Decube handles mixed-status selections intelligently. If some of your selected incidents are already in the target status, they're left untouched.
* If you're updating incident statuses to a new mute duration on already-muted incidents, the timer restarts based on the new duration from the moment you confirm changes.

**Traceability:**

* Each bulk action is logged in the incident's audit history with a `via bulk action` label, so your team always has full traceability.
* This can be found in the incident detail's page.

**Permissions:**

* Permissions are respected. Incidents you don't have edit access to will display a lock icon.
* These are automatically excluded from bulk selection.

See these changes in [Incidents Overview](/data-quality/data-quality#bulk-update-incident-status)

**\[Improvements]**

* **\[Snowflake Create/Manage Source Form]** When user tests connection, error messages now show up for the failure mode for incorrect input formats, wrong warehouse/accountname/private key/role
* **\[Asset Details > Schema tab]** List of columns are now alphabetically sorted.
* **\[Lineage]** Improved data fetching for lineages to handle case where a high number of lineages needs to be loaded.
* **\[Lineage]** List of columns of the entry point asset are now alpabetically sorted.

## \[Release Version: 1.51.0] 5 March 2026

**\[Feature Release: The New Decube Lineage Canvas]**

* **What’s new:** A high-performance canvas featuring fluid panning, mouse-wheel zooming, and a dedicated Minimap.
* **Contextual Pruning & Column Trace**
  * **Previously:** Users had to manually squint through a dense web of nodes to trace a single data attribute.
  * **Updated:** Selecting a column now automatically "prunes" the graph, hiding unrelated nodes to show only the relevant data flow.
  * **The Benefit:** Reduces visual noise by 90%, making PII audits and impact analysis significantly faster.
* **Smart Refocus & Viewport Stability**
  * **Previously:** Expanding a node often caused the layout to jump, pushing your active asset off-screen.
  * **Updated:** Introduced Active Pinning, which locks your focal node in place while the rest of the graph shifts to accommodate new data.
  * **New Tools:** Added Smart Refocus (one-click centering) and Instant Reset (clears all pathways) to keep your investigation focused.

See these changes in [Automated Lineage](/lineage/automated-lineage).

## \[Release Version: 1.50.0] 3 March 2026

**\[Feature Release: Sync Snowflake Tags]**

Decube eliminates manual overhead by automatically syncing Snowflake tags directly into the Data Catalog as System-Managed Attributes.

This integration ensures your catalog reflects the exact business vocabulary, security tags, and departmental classifications defined in your data warehouse without any manual configuration. Read more [here](/catalog/syncing-metadata-from-source/sync-snowflake-tags).

## \[Release Version: 1.48.0] 11 February 2026

**\[Feature Release: Asset Overview]**

The Asset Details page has been redesigned to optimize screen real estate and prioritize high-value data visualizations. Previously, the top header section occupied up to 40% of the vertical viewport when displaying attributes, pushing critical information like the Lineage graph, Incidents, and Monitors below the fold. The layout has been updated to be significantly narrower, providing a larger viewing area for the bottom tabs.

<figure><img src="/files/DfFEI49w7jAQAMPpiry8" alt=""><figcaption><p>Additional viewing height to explore complex lineages.</p></figcaption></figure>

Additionally, the "Overview" has been streamlined; all system-created and custom attributes are now consolidated into a single "All Attributes" section within the primary view. This ensures catalog consumers can access all curated information at a glance without navigating secondary tabs. For data stewards and owners requiring in-depth technical metadata (such as column counts or tables without owners), this information has been moved to a dedicated secondary tab to maintain a clean, focused workspace for general users.

<figure><img src="/files/HsR2PSyyAhMeGixJCp1A" alt=""><figcaption><p>Attributes all collated into one view on the Overview.</p></figcaption></figure>

**\[Enhanced Connector Support: Databricks Jobs & Pipelines]**

1. Added support for ingesting Databricks Jobs metadata. Ingested components include:

* **Data Jobs:** Definition, Last Run, Last Status, Schedule, and Job Runs.
* **Data Tasks:** Definition, Task Run, Last Run, and Last Status.
* **Supported Types:** Notebook, Python script, SQL query/file/alert, Run job, If/else conditions, For each loops, Python wheel, JAR, and Dashboards.

2. Implemented metadata ingestion for Databricks ETL Pipelines, including job definitions and run history.

* *Note:* This excludes Ingestion pipelines where the dataset originates from external sources.
* Lineage Attribution: Enhanced the lineage graph to show specific data job attribution for table-to-table relationships from Databricks Pipeline.

## \[Release Version: 1.47.0] 30 January 2026

**\[Deprecation of Data Quality Legacy Tests]**

We have officially deprecated all legacy, rigid Data Quality monitors and successfully migrated them to our new Unified Test framework. This migration consolidates separate test types (e.g., `Not Null` vs. `Null%`) into a single, flexible test with configurable thresholds.

## \[Release Version: 1.46.5] 21 January 2026

**\[Granular Change Request Tracking - Catalog Asset Details]**

We have overhauled the Change Request review workflow, introducing a granular, itemized summary of all modifications.

{% embed url="<https://www.loom.com/share/802aa1c401364a039c5e021c3d92e037>" %}

Previously, generic "Asset Updated" notifications forced reviewers to manually compare versions to identify changes. This update eliminates the "spot the difference" workflow, ensuring that every addition, removal, or edit is explicitly listed for faster, risk-free approvals.

This enhancement lays the foundation for advanced governance capabilities, with support for Business Glossary tracking and Visual Diffs following in the upcoming roadmap.

## \[Release Version: 1.46.0] 18 December 2025

**\[Data Quality: Unified Thresholds & Scorecards]**

We have introduced a unified configuration for Data Quality testing, streamlining how you monitor Null, Unique, Email, and UUID constraints. Users can now switch seamlessly between Row Count, Percentage, or Automatic thresholds within a single interface.

* Previously, legacy percentage-based tests (such as Null% or Unique%) were excluded from high-level reporting; these new Unified Tests are now fully integrated into your Data Quality Scorecard calculations, providing a more accurate and comprehensive view of your data health.
* Additionally, Custom SQL and Regex monitors have been upgraded to support these threshold configurations, allowing for more granular control over your custom validation logic.

<figure><img src="/files/ZPtNxg5Jo04GRZD6oC8N" alt=""><figcaption></figcaption></figure>

**\[Public API: Monitor Schema Update]**

To support the new threshold capabilities, the Public API has been updated with a revised schema structure.

* All existing monitor configurations fetched via GET requests will now return the new schema structure. This change ensures consistency between the UI and programmatic configurations.
* Existing tests utilizing thresholds have been automatically migrated to this new structure to maintain compatibility with the enhanced threshold UI.

See the changes [here](/public-api/overview/index/monitors/create-manage-and-delete).

**\[Governance: Policy Tag Enhancements]**

Previously, classification policy tags were restricted to a 25-character limit, which often forced users to use unclear abbreviations for complex compliance categories.

This limit has been increased, allowing users to create more descriptive and accurate policy tags that align better with organizational naming conventions.

## \[Release Version: 1.45.0] 2 December 2025

**\[Feature Release] Description Mode Selection (BETA)**

We’ve introduced a new Description Mode Selection at the source level, giving you control over how descriptions appear in the Catalog, either **User-managed** (editable in decube) or **Synced description** (read-only synced from source). This reduces clutter, improves clarity, and supports better documentation governance.

* Currently supported for **Snowflake** and **Redshift** sources.
* **Added support for schema object type:** Description Mode applies to Schema, Table, and Column levels, providing full consistency across all asset types.

**Key Highlights:**<br>

1. **Choose your mode:**
   1. While connecting or modifying a source (via *My Account → Integrations*), you can choose between user-managed or synced descriptions.
   2. Both options are stored in Decube, and you can switch modes anytime via Modify Source your selection applies instantly in the UI and CSV export.
2. **UI Updates:**
   1. Only one description field (based on selected mode) is displayed across Catalog.
   2. In Asset details -> Asset description field is moved to the asset attributes modal, you can edit and view table/schema descriptions in the asset attributes modal.
3. **CSV Enhancements**
   1. Default export includes the selected description (based on mode).
   2. New checkbox added to Export Catalog (Tables/Datasets) form to include both descriptions for auditing or comparison.
   3. Synced field appears as Source description (read-only) in CSV.
   4. Only user-managed description can be imported - remove source description before import.
4. **Public API**
   1. Both description and source\_description are returned via API for integration flexibility.

**\[Feature Release] Custom Attributes Export/Import**

You can now bulk update Custom Attributes across your Catalog and Glossary.

1. Export CSV: You can now see the Custom Attributes attached to an asset in the exported CSV.
2. Import CSV: To apply custom attributes to an asset, you can add columns in the CSV with headers `CA:<Attribute_name>` with their corresponding values to add them to the Catalog.
   1. You will be able to add custom attributes with its corresponding value to an asset, update its value or remove it entirely via the Import function.
   2. Importing operation requires specific format and rules. Review them in [https://github.com/DecubeIO/decube-docs/tree/public/csv-template-structure-edit-existing-items.md#custom-attribute-columns-ca-less-than-attribute-name-greater-than](https://github.com/DecubeIO/decube-docs/tree/public/csv-template-structure-edit-existing-items.md#custom-attribute-columns-ca-less-than-attribute-name-greater-than "mention")
   3. **Pre-requisite:** You will need to create the Custom Attributes first in the Org settings before proceeding to use them in the Import.

## \[Release Version: 1.42.0] 13 November 2025

**\[Custom Attributes: Search & Filter]**

* Adds an "Attributes" filter section to the Catalog sidebar that lets users build and apply advanced, org-level custom attribute filters.
* Users can add multiple attribute filter rows in a single modal session, combine rows with AND/OR, and see active attribute filters displayed as chips in the sidebar.

<div><figure><img src="/files/x7I4tGcSc3MbgYD0kBHn" alt=""><figcaption></figcaption></figure> <figure><img src="/files/UPfqcLPThrivx8QhcH4f" alt=""><figcaption></figcaption></figure></div>

**\[Custom Attributes: Public API]**

* New APIs for [Custom Attributes management](/public-api/overview/index/custom-attributes): GET, POST, PUT, DELETE operations
* Custom Attributes added to [GET, PATCH Assets APIs & POST Search Asset API](/public-api/overview/index/assets)

<figure><img src="/files/RpiFyWxmFefKWcWEKINb" alt=""><figcaption></figcaption></figure>

**\[User management: Invite User]**

* New: Assign groups during user invitation — Administrators can now select one or more groups in the Invite Users modal to automatically add invited users to the selected groups after they accept the invite.

## \[Release 1.41.0] 12 November 2025

**Data Mesh Module Deprecated**

* The current Data Mesh module, which included Data Domains, Data Products, and Data Contracts, has now been fully deprecated and removed from the platform.
* This deprecation marks the retirement of the existing Data Mesh experience to pave the way for a redesigned approach that will better align with core platform features such as ownership visibility, cross-linking, and configuration integration.
* All associated data and configurations have been removed, and documentation has been updated to reflect this change.

## \[Release 1.40.4] 24 October 2025

**\[Improvement]**

Enable multi-catalog ingestion on Databricks. Previously, one Databricks connection only allowed one catalog to be ingested. This has been updated so that all catalogs within a Databricks warehouse is ingested within a single connection.

* `Catalog` field has been removed from the connection form.
* Refer our Databricks connector documentation here: [Databricks](/warehouses/databricks)

## \[Release 1.40.0] 16 October 2025

**\[Feature Release] Migration of Data Sources to Integrations tab**

This update is designed to make managing your data connections more intuitive and efficient. We have moved **Data Sources** into the **Integrations** tab.

<figure><img src="/files/W0X79osM0FHGj81o603J" alt=""><figcaption></figcaption></figure>

This change centralizes all your data connection management into a single hub, making it easier for you to create and manage all integrations to your Decube account in one place. Here are the list of changes:

* Data sources now shown as cards (icon, status, name, owner) with View button.
* Search and filters on Data Sources list (search by name; filters: status, connector, owner) plus empty states and active-filter UI.
* Connector gallery to add new sources with static category tabs and search.

<figure><img src="/files/Mv8ink8EZI9HKOmPY8dN" alt=""><figcaption></figcaption></figure>

* Each connector page now has embedded documentation. This allows you to quickly reference supported feature capabilities and permission requirements without leaving the page.

<figure><img src="/files/Nzz1PA5HAqJkaLXu1the" alt=""><figcaption></figcaption></figure>

## \[Release 1.39.0] 13 October 2025

**\[Feature Release] Attribute Types: Integer, User, Enum, Boolean**

This release expands Custom Attributes beyond free-form text. You can now define organization-level attributes with the following types: Integer, User, Enum, and Boolean — and apply them to Catalog asset details and Glossary objects. Key highlights:

1. **New attribute types**

* Integer — numeric attributes with validation for whole numbers.
* User — link attribute to one or more users (single-select or multi-select depending on the attribute configuration).
* Enum — pre-defined option lists; editing an Enum value affects all existing attribute values that use that Enum.
* Boolean — simple true/false toggle for quick flags. See more details [here](/org-settings/custom-attributes)

2. **Where to use**

* Apply these attributes on Catalog Asset Details (Asset Overview and Schema) and all Glossary object types (Glossaries, Categories, Terms).

**Improvements**

1. Right click on any incident in the Incidents table to open it in a new tab.

## \[Release 1.37.0] 2 October 2025

**\[Feature Release] Sync descriptions from source \[BETA]**

The Catalog now automatically syncs descriptions from your data source tables and columns from data sources directly into the Decube Catalog on Snowflake and Redshift data sources.

This one-way sync (source → Decube) helps keep your cataloged assets aligned with the source-of-truth descriptions in your warehouse. Read more [here](/catalog/syncing-metadata-from-source/sync-from-source).

<figure><img src="/files/3wqOm7u8uUpPHNxE1dsd" alt=""><figcaption></figcaption></figure>

**\[Feature Release] Custom Attributes for Glossary**

This release adds Custom attributes application to the Glossary, allowing you to add these attributes to any Glossary objects. Read more [here](/glossary/apply-custom-attributes).

## \[Release 1.36.2] 29 September 2025

**\[Feature Release] Support for Delta Tables**

* Support has been added to catalog Delta tables on S3 and ADLS connectors. Read more on each documentation here:
  * [S3](/datalake/s3#example-5-delta-table-s3)
  * [ADLS](/datalake/azure-data-lake-storage-adls#example-5-delta-table-adls)

## \[Release 1.36.1] 25 September 2025

**\[Feature Enhancement] Lineage click-through**

* You can now click on the Asset name in the Lineage tab and get to its Asset Details in a new tab!

## \[Release 1.36] 24 September 2025

**\[Feature Release] Custom Attributes**

This release introduces the ability to define organization-specific, text-based attributes and apply them directly to your catalog assets. It’s the first step toward a more flexible, context-rich catalog experience. Key highlights:

1. **Custom Attributes Management**
   * Create, edit, and delete text-based attributes from Org Settings > Custom Attributes.
   * Choose object types and data sources for precise attribute targeting.
2. **Asset Details Updates**
   * Refreshed Overview tab layout with a dedicated Custom Attributes section.
   * Unified layout across object types to ensure Custom Attributes are displayed consistently.
3. **Custom Attributes Application**
   * Add and edit attributes directly on asset details pages, including schema/column level.
   * Integrated with change request submission and review workflows.

See more in the documentation [here](/org-settings/custom-attributes) for attributes management and [here](/catalog/data-catalog/apply-custom-attributes) for application.

## \[Release 1.35] 22 September 2025

**\[Feature Release] Support for Views**

* Views (which were shown as Tables indiscriminately previously) are now displayed with a distinct icon for Views and label for easier identification in the UI.

<figure><img src="/files/bI3cWAf5RwA53cFLuSbb" alt=""><figcaption></figcaption></figure>

**\[Feature Release] Data Quality Score & Dimensions for Custom SQL tests**

* Custom SQL tests now [allow the configuration of **Total Row Count Query**](/data-quality/how-to-set-up-monitors/custom-sql-monitors#dq-scorecard), which allows monitors to collect the row count for Data Quality scoring. When the Total Row Count Query is enabled for a test:
  * The calculated DQ scores populates the selected Dimension in the Dashboard and Reports.
  * Monitor History now show the **Total Row Count** for each scan when the **Total Row Count Query** is enabled in the configuration.
  * Reports will show the **Total Row Count** aggregated and latest count for scoring.
* Public APIs are updated to reflect the above changes to the Custom SQL tests.

<figure><img src="/files/YoIKCRuRVNF3eye2jKj1" alt=""><figcaption></figcaption></figure>

## \[Release 1.34.21] 17 September 2025

**\[Feature Release] Public API - Create, Modify & Delete Monitors**

* Users can now integrate with the Monitors API to create, modify and delete various monitor types on the platform.

**\[Feature Enhancement] Naming convention for Datalake Datasets**

* Improve the dataset naming convention for S3 and ADLS dataset to only show the final folder name in the pathspec as the dataset name.

## \[Release 1.34.18] 29 August 2025

**\[Feature Release] Feature Flag for Access Request Feature**

* Introduced a new setting to control the availability of Access Request feature to organizations.
* By default, no-access assets are shown with limited information so users can discover them and request access. With this feature switched off, users would not be able to discover assets they do not have access to.
* With this setting, organizations can choose to completely hide no-access assets, offering stricter visibility controls.

{% hint style="info" %}
If you would like to switch on/off the Access Request feature on your account, please reach out to your Account Manager.
{% endhint %}

## \[Release 1.34.17] 28 August 2025

**\[Feature Enhancement] Dynamically Truncating Glossary names**

* Glossary names now dynamically truncate according to your screen's aspect ratio, allowing you to view long glossary names in its entirety if you're on a larger screen size.

<figure><img src="/files/c4d1fR6fg55IJiVNKAxh" alt=""><figcaption></figcaption></figure>

## \[Release 1.34.14] 21 August 2025

**\[Public API] New Monitor Endpoints**

{% content-ref url="/pages/dqzq5MYrcKtFndI2eX83" %}
[Monitors](/public-api/overview/index/monitors)
{% endcontent-ref %}

1. GET Monitor
2. POST Search Monitors
3. POST Enable/Disable Monitor

## \[Release 1.34.13] 21 August 2025

**\[Feature Enhancement] Inclusion of Custom SQL into Reports**

* Custom SQL results are now available to be downloaded in the Reports module.
  * CustomSQL results are now added to the report generated, with `agg_error_row_count` included in the Reports.
  * Future sneak-peak: To allow calculation of DQ Score in the Reports and Dashboard, we would be including the ability to add a user-defined denominator in the CustomSQL. This will allow you to add a method to calculate the score from CustomSQL rules and assign a Quality dimension.
* Data Quality Reports generation now support Test Type filtering with Field Health and Custom SQL options.
* Custom SQL results are also added into the [Reports AP](/public-api/overview/index/reports/data-quality-scorecard)I.

<figure><img src="/files/DCgylpZEeue0YIkwVvff" alt=""><figcaption></figcaption></figure>

\
\&#xNAN;**\[Feature Enhancement] Config**

Filter with precision :rocket: — Config now supports Test Type filtering with drill-down sub-types, consistent with Incidents.

<figure><img src="/files/msiho6G1O8u21T8Sl4SR" alt=""><figcaption></figcaption></figure>

## \[Release 1.34.11] 18 August 2025

**\[Feature Enhancement] Data Governance**

We’ve leveled up Data Governance 🚀! Clearer navigation, in-tab rule search, improved feedback, and a refined review process: all designed to make workflows smoother and more reliable. Rule creation now supports grouped requests, auto-add for future assets, and distinct review modes for faster, more transparent approvals. **Here’s what’s new:**

1. **UI & Usability Enhancements**

   * **Simplified Navigation:** `Policy Info` and `Rules` are now in separate tabs for a cleaner, more focused experience

   ![](/files/j7aNp1AG1FmTjHyPULah) ![](/files/g6WwtMcL9zw23pqT6w3s)

   * **Rule Search:** Search directly within the Rules tab to quickly locate specific entries

   ![](/files/toi2DtM7uCcn7TlAin7A)

   * **Improved Feedback:** Tooltips clarify the difference between masking types (policy vs. rule). Clear toast messages now appear for all key actions — success, error, and updates.
2. **Rule Creation & Review Improvements**
   * **Grouped Change Requests:** Users can now bundle multiple changes (rule edits, policy updates) into a single request. The summary view under `Manage Requests` → Rules tab shows all changes categorized as Created, Modified, or Deleted — including asset-level differences.
   * **Reset Button:** A `reset` option has been added in the rule creation modal, with a tooltip warning that current selections will be cleared.
   * **Auto-Add for Future Assets:** Rules can now be created in advance with `auto-add` enabled. A tooltip clearly explains how new matching assets will be automatically included.
   * **Clear Review Flow:** Rules now support distinct `View`, `Modify`, and `Review` modes — with side-by-side change comparisons for easy approval.
   * **Rule Info Card:** An info card has been added to the `Already Selected` tab to provide clarity on what basis current assets were selected when the rule was originally created.

**\[Feature Enhancement] Glossary**

Introduced `Copy Link` action for glossary assets, allowing users to quickly copy and share direct links to a specific asset for easier reference and collaboration.

**\[Feature Enhancement] Public API - New Monitor Endpoints**

{% content-ref url="/pages/dqzq5MYrcKtFndI2eX83" %}
[Monitors](/public-api/overview/index/monitors)
{% endcontent-ref %}

1. GET Monitor History
2. POST Enable/Disable Scheduled Monitor
3. POST Trigger On-Demand Monitor

## \[Release 1.34.7] 8 August 2025

**Enhancements**

* Search functionality has been added to the Glossary. You can now directly search for Glossaries, Categories and Terms within the Glossary page itself.

<figure><img src="/files/MxsPCRdaMhX9OAebLM01" alt=""><figcaption></figcaption></figure>

* Catalog > Search results are also retained when you go back to it after moving to or accessing another page in the Catalog.

{% embed url="<https://www.loom.com/share/ee401d93876a40859906db88ee89d117>" %}

* Virtual table now has a new icon! This allows you to differentiate virtual tables from regular tables in the Catalog.

<figure><img src="/files/PAS6BtKRQ6FMPMENHHgn" alt=""><figcaption></figcaption></figure>

## \[Release 1.34.4] 1 August 2025

**\[Feature Release] Filtering by Classifications**

* Introduced the ability to filter both **Data Quality** and **Incidents** dashboards by **Classification Policies**.
* Enables users to view metrics and incidents only for datasets with specific classifications (e.g., PII, GDPR).
* Helps prioritize monitoring and issue resolution for sensitive or regulated data more effectively.

## \[Release 1.34.2] 30 July 2025

**\[API Enhancement] Data Quality Reports APIs**

We're excited to announce the release of our new Data Quality Scorecard API.\
This API allows you to programmatically retrieve, monitor, and report on the quality of your data assets. Now you can:\\

* Generate DQ Scorecard: Generate data quality scorecard reports for specified periods. Reports can be filtered by data source, schema, data quality dimensions, data owners, and tags.
* Retrieve DQ Scorecard Result: Access generated report results, including aggregated and latest DQ scores, error counts, and monitor details.

You can use the API to schedule a daily job to fetch scores, ingest the results into your data warehouse, and build a dashboard for proactive monitoring.\
Check out the documentation to get started: [Data Quality Scorecard](/public-api/overview/index/reports/data-quality-scorecard)

## \[Release 1.32.1] 28 July 2025

**\[API Enhancement] Assets APIs**

We've expanded our Public API with new Assets management endpoints, giving developers programmatic access to search, retrieve, and update catalog assets.

**New API endpoints:**

* **Search Assets:** Search and filter assets by type, parent, tags, classifications, and owners with paginated results
* **Get Asset:** Retrieve specific asset details by ID and type, including parent hierarchy and attributes
* **Update Asset:** Modify asset attributes such as descriptions, owners, and linked terms for various asset types

Supported asset types include sources, collections, datasets, properties, dashboards, charts, data jobs, policies, glossaries, categories, terms, and data domains. All endpoints support flexible filtering and comprehensive asset management capabilities.

For detailed API documentation and examples, see the [Assets API documentation](/public-api/overview/index/assets).

## \[Release 1.32.0] 10 July 2025

**\[New Module: Export/Import]**

We’re excited to announce the launch of Export-Import module—a powerful new way for data governance stewards and catalog users to manage metadata at scale. This feature is designed to streamline bulk updates, onboarding, and catalog maintenance, empowering your team to work faster and more efficiently.

**Key highlights:**

* Self-serve UI for instant export and import of catalog assets, glossary terms, and classification policies—no more waiting for support.
* Download ready-made CSV templates to add new glossary objects or classification policies, or export existing metadata for bulk review and editing.
* Import completed CSVs to apply updates or create new records in Decube, with smart validation and detailed error reporting.
* Supported objects include Catalog (datasets, dashboards, charts, data jobs), Glossary (glossaries, categories, terms), and Classification Policies.
* Strict access control: Only users in the Owner group can perform export/import operations, ensuring governance and security.
* All actions are logged in the History tab for full traceability, with downloadable error reports and operation statuses.
* Comprehensive documentation, best practices, and FAQ to help you get started and avoid common pitfalls.

For more details, see the [Export/Import documentation](/export-import/export-import-overview).

## \[Release 1.31.2] 20 June 2025

**\[Feature Release] S3 and ADLS Connector Upgrade**

We've updated the connector for ADLS and S3. This improvement allows users to:

1. Control the ingestion of files by defining it in the path specifications:
   1. **Control the format of CSV, JSON, JSONL:** You can now define the encoding, delimiter, if the file has headers and other settings in the path specification. This allows you to control how the collector will ingest the metadata of the datasets.
   2. **Include only datasets that match this pattern**: If the path spec specifies a table, and the regex is provided, only datasets that match the regex will be included. This gives you control over which datasets to be included into the catalog.
2. You can also see new metadata information collected from the dataset that is ingested into Decube:
   1. No. of Dataset Files
   2. Total Dataset Size

Check out the data sources page for the changes:

* [Amazon S3 Datalake](/datalake/s3)
* [Azure Data Lake Storage](/datalake/azure-data-lake-storage-adls)

## \[Release 1.30.0] 8 May 2025

**\[Feature Release] Application Principal Management System**

We’ve introduced a new authentication mechanism that enables identity-based access to AWS data sources using IAM roles and service identities.

* Allows customers to authorize Decube to connect to their AWS data sources using trusted IAM identities
* Supported data sources:
  * [Amazon S3 Datalake](https://docs.decube.io/datalake/s3)
  * [AWS Glue](https://docs.decube.io/transformation-tools/aws-glue)
  * [DBT Core with S3 storage](https://docs.decube.io/transformation-tools/dbt-core)

➡️ Learn more by clicking the links above for setup and configuration details.

*For more details on creating identities, refer to this* [*documentation*](https://docs.decube.io/security-and-connectivity/aws-identities)*.*

## \[Release 1.30.7] 2 May 2025

**\[Feature Release] OpenLineage Connector**

We have introduced a new OpenLineage connector to support ingestion of lineage metadata from third-party systems.

➡️ Learn more in the [*OpenLineage documentation*](https://docs.decube.io/transformation-tools/openlineage)*.*

## \[Release 1.29.0] 23 April 2025

**\[Feature Release] Added ability to move Term between Glossaries**

* New "Move Term" option added to the ellipsis menu in Glossary
* Users can select a target Glossary or Category from a dropdown list

*For more details, refer to the* [*Move Term documentation*](https://docs.decube.io/moving-terms-to-glossary-category)*.*

## \[Release 1.28.0] 31 March 2025

**\[Snowflake] Password authentication deprecation**

Snowflake password authentication has been deprecated. Our platform now only accepts Key Pair authentication to be inline with Snowflake's updated security policy.

## \[Release 1.27.0] 31 March 2025

**\[Feature Release] Multi Region Control Plane**

We have introduced multi-region control planes for Decube to enhance compliance and data residency. With this update, all production domains now operate on region-specific URLs.

**Changes:**

* Old Domain: <https://app.decube.io>
* New Regional Domains:
  * US Region: <https://us1.decube.io>
  * EU Region: <https://eu1.decube.io>
  * APAC Region: <https://apac.decube.io>

**Additional Notes:**

* Users are only allowed to log in to their designated regional control plane
* Cross-region logins will not be supported

## \[Release 1.26.0] 06 March 2025

* \[Feature Release] New Quality Dimension additions
  * Introduction of Timeliness, Consistency, Granularity and Others as new quality dimensions.
  * Monitors now have user-customizable Dimensions
  * Support for additional dimensions:
    * Generating Data Quality Reports
    * Dashboard > Quality subtab

## \[Release 1.25.0] 25 February 2025

* \[Feature release] Streamlined Config:
  * Can create multiple tests for each test type per column/table
  * Able to add a test name to differentiate monitors created.
  * Can create monitors in the Asset Details > Monitor tab directly without being redirected to another page.
  * Can modify monitors in Asset Details > Monitors tab.
  * Grouped by monitor is now additional setting in the normal test creation flow.
  * Each Schema Drift type is now separated by test type column, for consistency.

## \[Release 1.22.0] 23 Dec 2024

* Added Support for Oracle Integration: Decube now supports Oracle as a data source, enabling users to seamlessly create and manage Oracle connections. For details on setting up and granting privileges, refer to the [Oracle integration documentation](https://docs.decube.io/databases/oracle).

## Public API BETA- Phase 1

* We’ve launched Public API, available exclusively for Enterprise-tier customers. Platform Owners can now generate user-specific API keys through the Update Profile Settings under a new section below MFA. Users can generate multiple keys and access direct links to API documentation for supported endpoints.
* This release includes APIs for Glossary Management (create, update, delete terms) and Lineage Management (manual lineage definitions), providing seamless programmatic access to streamline workflows with Decube.
* For more details, refer to the [Public API Overview](https://docs.decube.io/public-api/overview).

## \[Release 1.20.0] 26 Nov 2024

Decube launched ADLS Monitoring, to enhance metadata tracking and event-based catalog management. This release supports Filter Modes such as Files Updated and All Records for more flexible monitoring workflows.

## \[Release 1.19.0] 08 Nov 2024

Decube is introducing a new AI feature that enables users to effortlessly generate asset descriptions in Catalog > Asset Details > Schema using Decube AI. This enhancement allows users to auto-generate comprehensive descriptions, improving metadata quality while reducing manual effort. The UI in Asset Details > Schema has been updated with an intuitive interface featuring a “Write with Decube AI” button for generating new descriptions, as well as a “Re-write with Decube AI” option for re-writing the existing ones, streamlining the process of adding and enhancing descriptions.

## \[Release 1.17.0] 14 Oct 2024

* **\[Iceberg with AWS Glue]:**
  * Decube introduced beta support for ingesting Apache Iceberg tables through AWS Glue Catalog. For more information on how to add this connector, you can refer to [Apache Iceberg (BETA)](https://docs.decube.io/datalake/apache-iceberg)

## \[Release 1.15.0] 03 Sept 2024

* **Azure Synapse Apache Spark Integration**\\
  * Added Support for Apache Spark Integration: Introduced a streamlined process for connecting Decube to Apache Spark within Azure Synapse Analytics Workspace. This includes the ability to create Spark data sources in Decube, integrate with Azure Synapse using OpenLineage, and configure Apache Spark for seamless data lineage tracking and management. For information, you can refer to [Apache Spark in Azure Synapse](https://docs.decube.io/transformation-tools/apache-spark/apache-spark-in-azure-synapse)

## \[Release Version 1.14.0] 02 Sept 2024

**Incidents Tab:**

* **Search Functionality**: Search functionality is improved now user can search for incidents based on asset:
  * Can search by direct asset (eg. search for column name to get incidents on columns)
  * Can search by parent asset (eg. search for table name to get incidents on the columns)
* Filtering Capability is added, now user can filter by:
  * Incident type & Test types
  * Incident levels
  * Incident status
  * Monitor mode
  * Assignee
* Schema drift incidents now show up in the Asset details - Incidents for the Asset.

## \[Release Version 1.11.0] 26 August 2024

**Glossary > Linked Assets**

* Improved user experience of adding a linked assets in the Linked Assets tab.
* Allow users to add new asset types as a linked asset: Dashboard, charts and more.
* Added new filters to the module which improves the searching functionality:
  * **Select Source:** Allows users to filter search results by data source selection.
  * **Filter by Asset Name:** Users can now search by various categories including Source, Collection, Dataset, Columns, Dashboard, Chart, and Data Job.
* Fixed issue where disabled sources were showing up allowing to be linked.

## 23 July 2024

**Config-Configure Monitors**

* UX is improved for setting up scheduled and on demand monitors. Incorrect user inputs (eg. thresholds) and required fields are now more clearly shown to inform users on which fields need to be filled before submitting the monitor configuration.

<figure><img src="/files/TzG8QFbLgkqqLUlfUouq" alt=""><figcaption><p>Overview for updated Set up scheduled monitor form</p></figcaption></figure>

<figure><img src="/files/MUP1dHZo8pYUL92omvwB" alt=""><figcaption><p>Overview for updated Set up on demand monitor form</p></figcaption></figure>

* **Custom SQL:** Under Custom Sql monitor, the character limit for the custom SQL monitor name is been increased to 100.

**Catalog-Asset Details**

* **Asset details monitors:** When clicking on configure monitors, user will be directed to the create/modify form for the monitor directly.
* **Export csv:** Now user can export the details of the table (description, tags, classifications) directly by exporting it to csv.

## June 2024

**Dashboard**

* New Dashboard is introduced with Overview, Incidents and Quality tab.

**Overview**

Dashboard provides an overview of:

* **Recently Accessed Assets:** Shows you the recent accessed assets from your dashboard.
* **Top Popular Assets:** Shows the recent activity of assets you have access to, including new data products, updates to the glossaries and others.
* **Recent Activity:** Shows the recent activity of assets you have access to, including new data products, updates to the glossaries and others.
* **Recent Requests:** Shows all access or change requests that have been raised for the assets you have ownership of.

**Incidents**

Under Incidents tab you get analytical incident overview with view of key incident metrics and summaries across domains and data sources.

* **All Incidents:** Shows the total no. of incidents in your organization data with the incidents status percentage for Open, Muted and Close.
* **Incidents Assigned:** Shows no. of Incidents assigned vs assigned.
* **Incident Levels:** Provides breakdown of all incident levels (Info, Warning and Critical).
* **Data Job Statuses:** Provides the data job status percentage (Pass, Fail and Others).
* **Data Contracts Breached:** Shows total no. of data contracts breached on data assets over a time period.
* **Average time to close incidents:** Shows total time required to close incidents shown across a time period.

Also provides a count of **total incidents** that had taken place over your selected timeframe. This information is also further classified into four types of data quality incidents:

* **Volume:** Triggers when there is a significant change in the size of data ingestion.
* **Freshness:** Triggers when a significant time has passed since the last update on the data.
* **Schema Drift:** Triggers when there is a schema change.
* **Field Health:** Triggers when anomalies are detected by our monitors based on the metrics configured by you. You can set up the monitors to track Field Health in the [Data Catalog](https://docs.decube.io/catalog/data-catalog).
* **Custom SQL**: Triggers when anomalies are detected by running the custom SQL.
* **Job Failure:** Triggers when anomalies are detected by data transformation, our system offers specialized Data Job/Job Failure monitors.

**Source/Domain Summary**

The Source/Domain Summary section provides a comprehensive view of all data quality incidents reported within a specified timeframe. Each source/domain will show the following metrics:

* **Tables With Incidents/ All Tables:** Shows the total no. of tables with incidents out of total tables.
* **Incident Count:** Shows the total incident count from assets in the data source.
* **Assigned:** Shows how many incidents have an assignee.
* **Closed:** Shows how many incidents with closed status.
* **Freshness:** This is the number of Freshness incidents.
* **Volume:** This is the number of Volume incidents.

**Quality**

The Quality tab provides a detailed view of data quality across various dimensions, supported by different test types. Here’s what you can expect:

**Quality Dimensions and Test Types**

Our platform currently supports four quality dimensions, each associated with specific test types:

* **Accuracy:** Measures how close the data values are to the true values. Tests include “Regex” and “Value in.”
* **Completeness:** Measures the extent to which all required data elements are present. Tests include “Not Null.”
* **Uniqueness:** Checks each data record to ensure it is unique within the dataset. Tests include “Is Unique.”
* **Validity:** Ensures data conforms to acceptable standards, such as ranges and formats. Tests include “Is Email” and “Is UUID.”

**Source/Domain Summary**

The Source/Domain Summary in the Quality tab provides results based on selected domains and shows scores for seven key quality metrics. This helps you gain a deeper understanding of your data’s health across different data sources and domains, making it easier to pinpoint areas for improvement.

## 21 March 2024

**Release 1.9.6 Updates**

**Linked Glossary/Business terms Column on Asset Details > Schema Tab:**

* Within the Asset Details > Schema tab, a new column titled "Linked Assets" has been added.
* Users can now easily see the linked glossary/business term associated with the asset.
* By selecting the linked term pill, users are directed straight to the corresponding glossary/business term, providing quick access to related information.

**Config Settings - Slack Notification configuration enhancement.**

* User are able to set up config settings just by submitting their Slack Channel ID or Slack channel name.

**Custom SQL - Grouped-by option is now supported !**

* This feature empowers users to organize their custom SQL queries by grouping desired columns and selecting distinct values according to their specific requirements.

## 22 February 2024

***Documentation***

Introducing our Documentation Feature, this feature allows users to add a knowledge base to an asset. Users can add notes directly to each data asset under asset details, facilitating seamless collaboration, knowledge sharing, and troubleshooting.

***Custom Frequency***

Our new Custom Frequency Option gives you complete control over how often your data is monitored. users are now allowed to set custom frequency based on the following frequencies:

* Daily: Users can select any hour of the day and a preferred time zone.
* Weekly: Users can select a day of the week, along with a preferable time zone and time.
* Monthly: Users can select between a monitor running on the last day of every month, or on a specified day of their choice, along with a time and time zone.

Simply select your frequency, set your credentials, and enjoy personalized data monitoring that fits your needs perfectly.

***Model Feedback***

Introducing our Model Feedback update, a user-friendly feature designed to fine-tune the level of sensitivity when providing feedback on incidents within your data monitoring platform. This tool helps users indicate how serious or urgent an incident is, making it easier for administrators to prioritize and respond accurately.

* **Adjustable Sensitivity Slider:** Upon rating an incident as "bad", using the thumbs down icon, users can easily slide the sensitivity scale to indicate the severity or urgency of the incident they're reporting.

***On-Demand Monitoring***

Welcome to our On-Demand Monitoring System, where you get instant insights into your data infrastructure whenever you need them. Here are the monitors that offer On-Demand Monitoring:

* Freshness
* Volume
* Custom SQL
* Field Health (except cardinality)

This feature allows users to create monitors to meet specific, one-off requirements.

## 29 January 2024

***What's New: Microsoft Teams Integration***

We are excited to announce the deployment of our latest update, introducing seamless integration with Microsoft Teams. Users can now utilise Microsoft Teams as their primary alerts medium and effortlessly configure custom alerts for their monitors through Decube.

**Key Features:**

* **Microsoft Teams Alerts:** Users can now choose Microsoft Teams as their preferred channel to receive alerts, enhancing real-time communication within teams. This can be configured at 'Config Alerts' under 'Config Settings' tab.
* **Custom Alert Configuration:** Tailor alerts for your configured monitors according to your specific monitor configuration needs.

## 18 January 2024

#### What's New

**Release 1.8.0 New Features:**

**Enhanced Asset Management:**

* Simplify permissions with the new "Request Access" feature, enabling users to conveniently request and grant permissions directly through the Catalog. This streamlined process ensures precise access for designated individuals.
* Designate ownership for connected data sources, bringing clarity and accountability to your system. Now, it's clear who is responsible for each connected data source.
* Flexibility in ownership extends to data sources, allowing seamless adaptation and reassignment of responsibilities within teams, fostering improved collaboration and resource management.
* Empower your teams with the ability to transfer ownership of schema and table assets, providing greater flexibility in adapting and reallocating asset responsibilities.

**Catalog Module Improvements:**

* Experience a more streamlined user interface with Catalog module enhancements, now including Schema assets and Data source assets. This update offers a comprehensive overview of tables within a schema and data sources, with the added ability to explore individual levels of selected data source assets.

**Asset Overview Update:**

* The updated Asset Overview in the Schemas and Data Sources catalog provides a more comprehensive summary for your schemas and data sources, including `counts of schemas`, `schemas without owners`, `total incidents`, `incident percentage`, `number of tables`, and `tables with incidents`.

**Access Request Management:**

* Enhance communication and streamline access requests with our updated "Request" tab within the Catalog module. Reviewers/Data Source/Schema/Data Source owners now receive a "Request Access" email, facilitating efficient management of access requests.

**Centralised Permissions Overview:**

* Simplify user access management with the centralized permissions overview in User Info and Group Management. Access a comprehensive overview of granted permissions on table level/group level requests, making the process more user-friendly.

**Lineage Revamp:**

* Explore our revamped Lineage feature, offering a dynamic and user-friendly interface to showcase your data flow. Experience dynamic upstream/downstream relationship lines and efficient column-level search functionality within the lineage table.

#### **Bug Fixes**

* Patched issues with the Redshift Connector Form.
* Patched issues with the Power BI Connector Form.

## 6 December 2023

**What's new**

**Single Sign-On (SSO)**

We are thrilled to announce the latest enhancement to our application, Single Sign-On (SSO) integration with Microsoft. This new feature is designed to streamline security and simplify the login process, providing a more seamless and secure experience for our users.

**Reports Module**

Introducing the latest "Reports" module on our platform. This powerful addition empowers users with valuable insights into dataset access and queries by user, enhancing transparency, accountability, and efficiency in managing data resources.

* Users can now generate detailed reports on dataset access, providing a comprehensive view of who has accessed specific datasets within the platform.
* The "Reports" module includes the ability to generate reports on dataset queries, showcasing the queries executed by individual users.

## 31 October 2023

**Release 1.7.3**

Added schema selection for Redshift and Synapse so users can only add the schemas that they want to add into decube.

## 30 October 2023

**Release 1.7.2**

Introducing the new Feed Feature in our Data Catalog module! This system on our platform allows users to interact, post threads, mention others, reply, and even edit their responses.

* Users will now get email notifications when they are mentioned, receive replies, or when comments are made on their data asset.
* Users are also able to subscribe or unsubscribe threads to stop/receive email notifications.
* Data assets now include a "Rating" feature, allowing users to evaluate them by giving ratings out of 5.
* A ratings filter was added to the Data Catalog, allowing users to sort data assets by their given star ratings.

**We also implemented new connectors!**

We added a new connector:

* AWS S3 Datalake.

## 2 October 2023

#### Release 1.7.1

Multifactor Authentication with authenticator apps is now supported.

* MFA can now be enforced in your organization by an Owner or a user with `Manage users for access management`.
* Users can set up MFA on their account to ensure security for the account.

## 26 September 2023

**Release 1.7.0**

New module - Data Governance. One place to manage classification policies in your organization.\
Users are able to store all governance policies in the platform to be managed as well as link actual data assets from the Catalog to the policy.\\

* Classifications are now customisable - users can now create custom classifications by adding a custom policy in the Classification Policies tab.
* Users can now use our auto-classification workflow to set a rule in connected data sources to automatically tag columns that match a keyword or regex input. Example: classify all first\_name columns with PII.
* All classified assets under a policy (eg. list of columns classified as PII) shows up in the Managed assets of the policy. Users can export the entire list in csv format.

Module-based policy in Group Management has been updated with new selection in dropdown that is Data Governance. Users will need to give access to the Data governance module by adding the required access through the module-based policy form.Revamped Config - Centralised control to create, update all monitors.\
Our config module landing page simplifies the process of setting up various monitoring methods for your assets. You can navigate here by the Config panel in the sidebar.\\

* All monitors in Config now shows all available monitors that your organization has added, for easier modification and enable/disable. Job Failures, Schema Drift, Custom SQL monitors all are now shown in one page based on your selected data source.
* Incident Levels (Info / Warning / Critical): You can now set individual incident details
* For Field Health, bulk monitoring is now supported. Instead of adding a single test to each column at a time, this has been improved to add tests to multiple columns at once.
* You can control Schema Drift monitoring in a more granular manner by controlling which schema change type to monitor, and which schema to be monitored.
* Smart training & Lookback period: You can now shorten the period of training for each monitor by enabling smart training. Adding a lookback period also allows you to backfill your monitor with historical data, so you can see any anomalous data from before the time the monitor was set up.
* Custom alert channels: You can now set up specific alert channels for each monitor. For example, link Airflow Job failures monitoring to a specific Slack channel #job-failures. This setting will override the default alert setting in your organization.
* With the addition of custom alerting, we have upgraded the notification system to send out individual alerts instead of grouping the alerts based on test type. With this change to individual alerts, each alerts will now contain more information for each incident, and clicking on the provided link will send you straight to the specific incident details page instead of landing on the Data Quality page, saving you time.
* Row Creation selection for monitoring: “Metric Time” has now been renamed to “Row Creation” for more clarity. There are several important changes here:
* All Records: Adding a monitor to scan the entire table is now supported.
* SQL Expression: If there’s no datetime/date column in your table to be selected as a timestamp column (eg. you store it as a string type column), you can now convert it by adding a valid SQL expression to be used as a timestamp column on our platform.

## 9 August 2023

#### What’s new

**Incidents Details**

* More informative graphs for each incident type, with new chart types and added axis titles.
* The ability to assign the incident to one of the users.
* See the owners for the affected asset in the incident details on the side panel.
* An incident history where user actions of changing incident status of muting, closing the incident is tracked.
* For Field Health incidents, a sample of the affected rows can be found directly in the Preview tab, where dynamic masking of PII information is applied.
* See all scans from your monitors with datetime and actual with expected values in the History tab. Users can click into any failed incidents to see the specific incident description.
* See all assets impacted downstream of the asset with incident in the Impacted Areas tab, users can also download the list as a csv.

### **8 August 2023**

#### What’s new

We added a new connector: AWS Glue. You can now see your transformations as a Data Job in the Catalog.

### **2nd August 2023**

#### **Bug Fixes**

* UPS error ("invalid asset") is resolved during group creation without any sourced-based policy added.

### **1st August 2023**

#### **Bug Fixes**

* Missing data sources during Sourced-Based Policy creation now appears.
* Clicking on back button on policy forms sends user back to Policy selection form instead of closing the modal.
* Source Modal Policies creation using template (Use as template) fixed.
* When user wants to add more than 3 tags, there is now an error prompt that is shown. Users will also be shown an error when attempting to create a tag with special characters.
* "Database" word removed from forms with non-database connectors.

## July 2023

### 31 July 2023

#### What’s new

* Search in our Catalog is now powered by Elasticsearch, allowing for a better search experience with natural language search. Searching by path has also been enabled, such as searching `schema_name.table_name`.
* Clicking on specific search suggestions from your search input now leads you directly to the asset. Pressing on “Enter” or clicking on the Search button searches your Catalog with your search input.
* Narrow down your search by quickly filtering assets by type, schema, tags, classifications, owners, and count of incidents, ensuring you find exactly what you need with minimal effort.
* Introduced a sort option to sort by relevance, or alphabetical order.
* Glossary items such as Glossary, Categories and Terms are also shown in search results. Clicking on these will bring you to the location of it in the Glossary directly.

#### Bug Fixes

* Classification not showing in the Glossary Terms has been fixed.

### 13 July 2023

#### Bug Fixes

Page crash experienced on Asset Details > Preview for Table assets has been fixed.

### 11 July 2023

#### Bug Fixes

Added patches for security vulnerabilities identified.

### 10 July 2023

#### Bug Fixes

* Fixed the issue of the password field on the organization invite form not allowing special characters. Added a list of special characters in the UI.
* The issue of schema selection in BigQuery form not being able to be submitted has been fixed.

## June 2023

### 21 June 2023

#### What's new

1. Our navbar became thinner! This gives you more space for your Lineage, Incident Details to take up more real estate on your screen.
2. Our lineage also now can collapse once you've expanded it, by clicking on the "-" sign.

**Bug fix**

Some of our tooltips were hard to read because of the sizing, so it is not fixed to fit perfect on each component.

### 16 June 2023

**What's new**

We've got a new connector - Azure Data Factory. Check out how to get connected [here](/transformation-tools/azure-data-factory).

### 14 June 2023

**What's new**

1. Group Management: Manage user access by creating groups and granting these groups permissions. Check out the summary [here](https://github.com/DecubeIO/decube-docs/tree/public/overview/broken-reference/README.md).
   * If you're an existing customer, we will migrate you seamlessly to the new system. Admins are now in the Owners group, and Members will be transitioned to the Members (Legacy) group. All data sources access will be remained. [Read more](/group-access-policies/groups-management-overview#transition-from-pre-gbac-to-gbac).
   * New customers who is the first user sign up will be added to the Owners group. [Read more](/group-access-policies/owners).
   * A guide has been made for create groups and assigning policies [here](/group-access-policies/create-groups-and-assign-policies).
2. Approval Workflow: Now you need to go through an approval process to make changes to the Catalog or Glossary. Check it out [here](/approval-workflow/summary-of-approval-workflow).

### 2 June 2023

#### What's new

1. New connector: Azure Synapse Analytics! We are continuously adding integrations to Azure products. Check out our [Public Roadmap here](https://decube.notion.site/aad545f86d9b4708a92fd0514adddff8?v=4b8e373745924bd687fb7bbc90002b26\&pvs=4).
2. Made our Schema Drift incidents to show description of the event that happened, eg. `addition`, `table`, or `type_change`.

## May 2023

### 26 May 2023

#### What's new

1. New alert type: [Webhook integration](/alert-notifications/webhooks-integration) so that you can set up your own custom integrations.
2. Upgraded our Lineage with our SQL parsing engine, you may see some new lineage relationships shown in the Lineage tab of your Asset details.
3. You can now disconnect your Slack organization on the Configure Alerts page.

#### Bug fixes

1. Issue where the decube Catalog does not show the right assets from your added Databricks Catalog is now fixed.

### 19 May 2023

#### What's new

1. Data Recon UI has been completely refreshed! This includes:
   * Adding new recon is now a more focused experience, taking up a whole page instead of being a sidebar.
   * You can compare datasets from two separate data sources.
   * Set up conditions to filter out rows to be checked in the recon. This functions as a WHERE clause. Read more [here](https://docs.decube.io/data-reconciliation/data-recon).
   * Selection of datetime field is now optional.
   * You can now export your completed recon in .csv. This downloads the unmatched rows found during the data recon.
2. Preview tab has been added for Asset Details.
   * We show a sample of 10 rows from your table.
   * If there are columns that are marked as PII or Sensitive, this will show up masked in the Preview.

#### Bug fixes

1. UI clashing issue with the sidebar on small window sizes has been fixed.
2. Fonts in graphs and charts have been updated to Plus Jakarta Sans.
3. Notify toggle on Modify monitor modal has been fixed. It now shows the correct state, whether it's enabled or disabled.

### 11 May 2023

#### What's new

New look on our app! We've updated our font for better readability across all our pages.

#### Bug fixes

1. Lineage graph not re-appearing when navigating tabs has been fixed.
2. When viewed on small screen sizes, the UI on Table Overview had clashing issues, which has been fixed. Screen sizes down to 720p on normal scaling are now supported without issues.

## April 2023

### 27 April 2023

#### Bug fixes

1. On Safari 15.6, page crash when viewing the field health monitors tab has been fixed. (Safari 16.1 and other browsers are not affected by this bug).
2. Job failure monitors: Toggle for Notification on Data Job monitors now work as intended.

### 20 April 2023

#### What's new

* You can now see the lineage from your transformations in your Lineage now! Just go to the Asset Details and head to the Lineage tab. You'll need to map your Lineage via the Additional Config tab in the Modify Data Source under My Account.
* Support for Azure SQL has been added.

#### Bug fixes

* Documentation editor icons have been fixed. The editor also looks better on wider screens now.
* Data Recon now shows the empty state when there are no new data recons added yet.

### 19 April 2023

#### Bug fix

Clicking on a table in Dashboard was sending users to a broken link. It's not been fixed and should lead to the Asset Details now.

### 13 April 2023

#### What's new

* Added new integration: SQL Server! You can now connect your SQL Server to decube to monitor your data health and see all your tables in the catalog. Read more about it [here](https://www.decube.io/post/new-connection-microsoft-sql-server).

### 12 April 2023

#### What's new

* Added a new integration: Singlestore! We're also have a partnership now. Read more about it [here](https://www.decube.io/post/maximize-the-value-of-data-with-singlestore-and-decubes-partnership).

### 5 April 2023

#### What's new

* Business glossary: Introducing a new feature where users can add their own metrics, add owners and documentation. You can also link them to data assets. The business glossary will allow users to create a common language for metrics and data assets, improving communication and collaboration across teams.
* New integration with Fivetran
* New integration with Databricks

## March 2023

### 21 March 2023

#### What's new

* New integration with **dbt**, supporting new asset types such as data jobs, data runs, virtual tables etc.
* New incident type, job failures has been introduced.
* Schema tab added, users can now add tags and classifications to fields.
* Field stats added, users can now query their datasets and find out some statistics on each field. Also integrated the data quality incidents into this view.
* A new look to the Data Quality and Dashboard page, with new asset icons.

#### Bug fixes

* Fixed a bug where Intercom is still recognizing the user when user is logged out automatically.
* User is now directly correctly to the filtered incident types in the Data Quality page from notifications.
* Incident graph in Data Quality main page are now bar charts!

### 9 March 2023

#### What's new

* For Google Big Query and Snowflake connections, you can now add Volume and Freshness monitoring **without** a metric time selection.
* For all sources, you can now use `date` column as a metric time column. Adding a `date` type column would restrict your scan frequency selection to `24 hours` only.
* Monitoring for schema changes is now enabled on all tables by default. You will just need to let us know if you'd like to receive notifications by toggling the notification within the Asset Details.
* UI improvements on the Bulk Table Config and Asset Table Config for better user experience, such as updated icons for metric time selection and descriptive placeholder texts. Also added clearer tooltip for users to be aware of the pre-requisite to select the metric time before a field monitor can be added.

#### Bug fixes

* Lots of bug fixes and improvements made on the Data Quality engine for better reliablity and flexibility.

### 6 March 2023

#### Security

If user becomes inactive, after some time the cookie will expire and user will be redirected to the log-in page. Logged in user who is actively using the application will have user session automatically extended.

### 2 March 2023

**What’s New**

* Sign-up and sign-in pages have been updated to be more responsive and make the sign-in process smoother.
* We are now supporting your organization via Intercom! Drop us a message via the bubble on the bottom right if you need any help.
* The UI for adding and modifying monitors has been updated.

#### Bug Fix

Fixed new data recon: Scheduled run period not submitting when user did not interact with the dropdown.

## February 2023

### 23 Feb 2023

**What’s New**

* For Google Big Query connections, you can now specify which schemas are to be scanned into decube.
  * If you previously had connected GBQ, you can now go to My Account and modify the connection and unselect the schemas that you would not like to include. These schemas will be removed on our next collector run.
  * You can also opt to not automatically add new schemas in the future when they are detected by our collector by selecting your preferred option.
* Don’t like the name that you used while signing up? You can now update it via the Update Profile tab on My Account.
* You are now able to add Owners (data owner and business owner) to an asset. For now, tags on tables will be unavailable as we are improving this functionality to add the tags directly on the field level such as PII, Sensitive tagging which is coming very soon.
* You can now save time by updating your table-level tests (Freshness, Volume, Schema drift) from the Data Quality page directly.
* The incident details graph has been improved to show the min/max of the Expected vs the Actual for Volume graphs.
* Our docs has been updated for easier navigation to specific connections and features.

#### Bug Fixes

* Fixed the count of data sources on filter selection when data sources were disabled.
* Fixed behavior of custom monitors lists so that it does not show lists of other tables.

### 8 Feb 2023

**What’s New**

User will need to accept the disclaimer on the Table Overview (which is an opt-in only feature) before the profiler is ran.

**Bug Fixes**

* Bug fixes on the UI on the Dashboard and the bulk selection.
* Terms and Privacy Policy on the sign up page have been updated with new links.

### 1 Feb 2023

#### Bug Fixes

There is now a sorting logic on the Bulk Table Configuration so that monitored tables appear on the top of the list followed by selections with metric time. This helps you to be able to see configure your tables easily.

## January 2023

### 31 Jan 2023

#### What’s New

* Plans & Billings page under My Account now shows the usage limit according to the plan the organization is subscribed to.

#### Bug Fixes

* UI fixes across the app such as truncating long asset names and rephrasing for a better user experience.

### 30 Jan 2023

#### Bug Fix

Creating a new data recon now shows loader when operation is running. If a recon job is not successful, a failure state will be shown.

### 25 Jan 2023

#### What’s New

A new connector to Snowflake has been added. Users can now add Snowflake as a new data source under the My Account page.

#### Bug Fixes

* When creating a new Data Recon, you can now search for your tables that you would like to use.
* When there are no incidents yet to be shown on the dashboard, we now have a more descriptive text on the UI.

### 20 Jan 2023

#### What’s New

The Lineage UI has been enhanced with a cleaner look and new icons.

#### Bug Fixes

* Disabled data sources now do not show inside data sources selection.
* Incidents list within Table details have been fixed.

### 19 Jan 2023

#### Bug fixes

Fixed issue in Bulk Table Configuration where selections were not able to be viewed when there are many metric time selections.

### 16 Jan 2023

#### What’s New

* Catalog has been enhanced with new icons and filtering for search functionality.
* User can now add new monitor directly from the Data Quality page.
* Get support via live chat on the bottom right of our page.
* Custom SQL monitors can now be added:
  * User will be able to add Custom SQL scripts to write specific tests cases as monitors.
  * Monitors page within Table details has been updated to support custom monitors which has user-defined names and descriptions.

#### Bug fixes

* Fixed issue where more tables and incidents were not loaded when scrolling to the bottom of the page in the Dashboard and Data Quality.
* Fixed issues on Data Recon page.

### 10 Jan 2023

#### What’s New

* User are now able to use Bulk Table config feature to do bulk selection on multiple metric time columns. This new feature can be accessed via the Catalog.
* Users will be required to configure their tables’ metric time columns and monitoring before enabling table-level and field-level tests. Switching on monitoring for a table enables all table-level tests to be on by default.
* Schema/dataset name has been added to new data recon selection.

#### Bug fixes

* Incident debug messages have been improved.
* Adding tags to assets, schema drift and table config overlay issue with metric time selection have been fixed.
* Removed error logs on console to prevent info leaks on logs.

### 9 Jan 2023

#### Bug fixes

* Ensure users are not able to log in after deactivation of account.

### 6 Jan 2023

#### What’s New

* User management
  * Admin able to change user roles of other users (`member` to `admin` and vice versa).
  * Admin able to revoke invites of users who are `pending sign up`.
  * Admin able to deactivate/re-activate other users.

### 5 Jan 2023

#### Bug fixes

Safari users facing issues with UI freezes on certain pages with loading functionality which was fixed with his hotfix.


# Support

Here is how you can reach us for support enquiries.

If you have any questions or need assistance, feel free to reach out to us at <mark style="color:blue;"><support@decube.io></mark>. We're here to help!

During onboarding or ongoing collaboration, we can also set up a dedicated **Slack** or **Microsoft Teams channel** for more direct support. Just let us know via your assigned Account Manager if you'd prefer this, and we’ll get it arranged for you.

### Quick links

{% content-ref url="/pages/cmLNwa8CruqVyQNcltOP" %}
[Frequently Asked Questions (FAQ)](/overview/support/frequently-asked-questions-faq)
{% endcontent-ref %}

{% content-ref url="/pages/S9LppF7GJ4s95RF28t25" %}
[Broken mention](broken://pages/S9LppF7GJ4s95RF28t25)
{% endcontent-ref %}

{% content-ref url="/pages/m9CuNY35z7hByvbK3qWF" %}
[Supported Browsers and System Requirements](/overview/support/supported-browsers-and-system-requirements)
{% endcontent-ref %}


# Frequently Asked Questions (FAQ)

Find quick answers to common questions about our platform, features, and support. If you can’t find what you’re looking for, feel free to reach out to our support team.

## Data Quality

**Why isn’t my monitor triggering any incidents?**\
Check if any tests are assigned to the monitor. Monitors need at least one failing test to trigger an incident.

**Which monitors provide an incident preview when generating incidents?**

* Only **Scheduled Field Health** monitors provide an incident preview.
* **On-demand** monitors do **not** have an incident preview, as they are designed for one-time runs based on user configuration.
  * Firstly, Decube **does not store any raw data** for Preview. All previews run are a one-time query to your data.
  * When setting up an **On-demand test**, the user sets a lookback period, which will run a lookback of x hours/days of data. This means that when you set a lookback period of 1 hour, the monitor will run a check of the data of the past 1 hour and raise an incident if there are any.
  * After time has passed, the on-demand test incident preview may not provide an accurate representation of the data that triggered the incident failure.

***

## Catalog

**What does each connector support?**\
Each connector documentation has a list of supported capabilities at the top of the page of the connector documentation.

**How often does Decube refresh the data catalog?**

Decube refreshes the Catalog on an hourly basis. Throughout the hour, we perform automated metadata ingestion on your data sources to keep your Catalog updated.

**Why are some of my columns not appearing in the Field Statistics?**

The current version of our profiler does not support the profiling of date, time and datetime columns. If this is something you wish to see in the future, please let us know via the Live chat.

***

## Config

**What are the differences between enabled and disabled Smart Training?**

When **Smart Training is enabled** (e.g., for Row Creation):

* The monitor immediately scans **historical data** based on available records.
* It uses the **confidence level** to determine an appropriate threshold.
* This allows the test to be **created and run quickly**, often without needing to wait for multiple executions.

When **Smart Training is disabled**:

* The monitor must run **a few times** to collect enough data to establish a baseline.
* This can take **several days**, as it doesn’t scan historical data retrospectively.
* During this period, the test may be marked as **"Skipped"** until enough data is available.

<i class="fa-circle-info">:circle-info:</i> We recommend enabling **Smart Training** to accelerate the learning process and reduce the time needed to generate meaningful test results.

**Why is the test status marked as Skipped if Smart Training is not enabled?**

Even though the **monitor is running**, the test status may appear as **"Skipped"** because there is **no historical data** collected yet. Without **Smart Training** enabled, the system requires a few days to gather enough historical data before it can generate and execute the test. Until sufficient data is available, the test cannot be performed, so Decube marks it as **"Skipped"**.

**What does the Freshness Monitor actually do?**

The **Freshness Monitor** is designed to track **how frequently a table is updated**, not just when it was first created.

**Does Freshness Monitor only check based on the day-1 of the row creation column? What if our table is updated multiple times a day?**

If your table is updated multiple times a day (e.g., hourly), the Freshness Monitor will compare the **last scan time** with the **most recent update timestamp** of the table. This allows you to detect if data is being refreshed **within the expected frequency**.

**Can we check same-day freshness using a scheduled monitor?**

**Yes**, you can monitor **same-day freshness** by configuring the monitor to run on a **scheduled basis** (e.g., hourly, every few hours). The monitor will validate if the table has been updated within the defined freshness threshold.

**How do the monitors for Schema Drift and Job Failures work?**

**Schema Drift**

* Decube begins by **ingesting metadata** on a scheduled basis from connected data sources.
* It then **compares the latest schema** from the current scan with the schema from the previous scan.
* If any structural changes are detected — such as **columns being added, removed, renamed, or data types being changed** — a **Schema Drift incident** is automatically created.

**Job Execution**

* Decube monitors data job executions by collecting metadata that includes the run results, whether a job has passed or failed.
* For example, Decube ingests the **dbt Core manifest** stored in **S3**, then uses this manifest along with job metadata to build **data lineage**. During each scheduled metadata ingestion (e.g., hourly), Decube reads the **job run status** to check for any failures.
* This enables teams to proactively monitor their pipelines and get notified when a job fails, helping to maintain the reliability of data operations.

**In the Data Catalog, the default sorting order is Relevance. How is this calculated? Can it be changed to A-Z?**

The **Relevance** sorting is based on several search factors, including how well the search term matches metadata fields such as the asset’s **name**, **description**, and more.\
This is powered by the **Elasticsearch** engine, a widely used industry-standard search technology.

We default to **Relevance** to ensure more accurate and useful search results.\
If we defaulted to **A-Z**, the top results would be sorted alphabetically, not by how closely they match your search query — making it harder to find what you're actually looking for.

Currently, it's **not possible to change the default sorting order**, but you can manually sort by A-Z after performing a search.

***

## Data Governance

Is PII data masked in Decube?

Yes, once you have applied a classification to these PII data, it will be **masked** in Decube when user is viewing the Preview through the Catalog.


# Supported Browsers and System Requirements

decube is a web based application that does not require any download or install for use on a computer with a supported web browser.

Our application is rigorously tested and supported on the following web browsers on desktop:

* Chrome
* Edge (Chromium based only)
* Firefox
* Safari

The latest 2 versions of each browser in the list above are supported.

{% hint style="warning" %}
Internet Explorer (IE) is not supported.
{% endhint %}

Currently, our application is not supported on mobile web browsers; it's recommended to always use the desktop version to log in.


# Overview

We built our unified data platform with industry standards to protect your data and ensure compliance while delivering observability and governance.

Decube is designed from the ground up to ensure that enterprise-grade security, compliance, and scalability are built into every layer of our platform. This document outlines Decube’s infrastructure design, security protocols, and compliance posture.

For deployment-specific models details, refer to our [Architecture ](/security-and-infrastructure/deployment-methods).

## How we handle your data?

### Control & Data Plane Separation

Decube follows a strict separation of concerns between the Control Plane and the Data Plane:

* **Control Plane (Decube-managed)**: Handles authentication, user management, licensing, configuration, scheduling, and alerting.
* **Data Plane (deployment-dependent)**: Executes metadata collection, monitors data quality, and stores metadata, either in a shared or isolated environment based on the deployment model.

Details for each model are available in:

* [Multi-Tenant SaaS](/security-and-infrastructure/deployment-methods/saas-multi-tenant)
* [Single-Tenant SaaS](/security-and-infrastructure/deployment-methods/saas-single-tenant)
* [BYOC](/security-and-infrastructure/deployment-methods/bring-your-own-cloud-byoc)

### Collection

* Decube's data collectors only extract metadata, query logs, and aggregated statistics into its cloud service.
* Data extracted from these scans is solely for assessing your data's reliability and providing statistics and incident alerts of which you have opted-in.
* Decube uses encrypted connections (HTTPS and TLS) to protect the contents of data in transit.
* Decube's architecture also supports a setup specifically for enterprise customers where you can host the data collectors within your own cloud infrastructure so you never have to expose any of your data sources to decube's cloud service.

<figure><img src="/files/3wDZD7Mtpy3gNXhGDBtV" alt=""><figcaption></figcaption></figure>

### Compliance

* Decube is currently SOC2 certified. Reach out to us if you would like a copy of the SOC2 report.
* Decube will sign any NDAs and/or DPAs where it is appropriate.
* Decube, while collecting metadata, query logs, and metrics for the purposes of running the monitoring, cataloging, acknowledges that personal data may be collected and processed. If any such data is passed into Decube, it is used only for the sole purpose of running the monitoring and cataloging.
* Usage of all SaaS applications internally within Decube for operational purposes is vetted with due diligence so that confidential company and personnel data are protected.

### Organizational Security and Privacy Practices

Decube's team practices industry best practices across the board to protect the security of the application, and the data privacy of its customers.

* Decube engages a third party to perform an annual penetration test over the application layers of the platform.
* Processing of collected data is conducted on secure servers hosted on Amazon Web Services.
* Decube employees engage in privacy and security training during the onboarding and are required to take an examination after the training. All Decube personnel are required to acknowledge, electronically, that they have attended training and understand the security policy.
* Access to all critical systems and production environments are protected using strong passwords and multi-factor authentication. SSO is also used to centralize access control for certain applications. Access rights are reviewed before being granted, and then periodically reviewed thereafter.

### Data collected by Decube

The following information may be processed and stored by Decube on its cloud services:

| Collected Data        | Details                                                                                                                                                 | Purpose of collection                                                                                                                                              |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Metadata              | Asset names such as tables and columns, field types, status of transformation jobs and other such metadata.                                             | To populate the data catalog with information about the assets available (table, columns, jobs etc.) within the data warehouses, databases and other data sources. |
| Metrics               | Row counts, last updated and other similar metrics.                                                                                                     | Enable tracking of metrics such as freshness, volume and other metrics.                                                                                            |
| Aggregated Statistics | Measures the data in selected table which is opt-in only by the user. Statistics may include null percentiles, distinctness, and other similar metrics. | Enable tracking of data health and setup of monitors by user via preset or custom SQL.                                                                             |

#### Multi-AZ Compute

Decube’s cloud infrastructure spans multiple AWS Availability Zones to ensure high availability and business continuity.

### Metadata Collection Approach

Decube collects only the metadata required for observability, governance, and quality monitoring:

#### Metadata Types Collected

| **Aspect**                  | **Description**                                                  | **Sample**                                                               | **Encrypted**         | **Retention** | **Location**                |
| --------------------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------ | --------------------- | ------------- | --------------------------- |
| **Data Source Schema**      | Schema metadata about data sources, broken into logical grouping | Table name, Column names, Column data type, Constraints, Dashboard names | At-Rest Encryption    | Indefinite    | Metadata DB & Elasticsearch |
| **Monitoring Metrics**      | Aggregated metrics from data quality monitoring                  | Row count, Nulls, Unique values                                          | At-Rest Encryption    | Indefinite    | Metadata DB                 |
| **Change Requests**         | User changes on metadata                                         | Asset description, Classification, Custom attributes                     | At-Rest Encryption    | Indefinite    | Metadata DB                 |
| **Profiling Info**          | Statistical values (may contain raw data)                        | Null %, Min/Max/Average values                                           | Dual Layer Encryption | 30 Days       | Blob Storage                |
| **User Attachments**        | Uploaded files                                                   | PDF, CSV, Text                                                           | Dual Layer Encryption | Indefinite    | Blob Storage                |
| **Query History**           | Historic queries on platforms                                    | Synapse, Databricks logs                                                 | At-Rest Encryption    | 60 Days       | Metadata DB                 |
| **Data Source Credentials** | Connection credentials                                           | Username, Host, Password                                                 | Dual Layer Encryption | Indefinite    | Metadata DB                 |

## Data Access & Encryption

All connections to customer systems use read-only credentials and are encrypted with TLS.

Decube supports dual-layer encryption:

* Content Encryption using AES-256-GCM
* At-Rest Encryption for data stored in our services

**Encryption Details**

* **Dual-Layer Encryption for Maximum Security**
  * Decube employs a robust **dual-layer encryption** approach to ensure the confidentiality and integrity of sensitive data across all deployment models.
* **Content Encryption - AES-256-GCM**
  * All sensitive data—such as query history, profiling metrics, and credentials—is encrypted using the Advanced Encryption Standard (AES) with 256-bit keys in Galois/Counter Mode (GCM). This method guarantees both data confidentiality and integrity by protecting the actual content from unauthorized access or tampering.
* **At-Rest Encryption - AES-256-GCM**
  * In addition to content-level encryption, all stored data benefits from at-rest encryption. This additional layer secures data while it’s stored in databases or object storage, offering comprehensive protection against breaches or unauthorized access at the storage level.
* **Indefinite Retention**
  * Data stored until the underlying data source is deleted from Decube’s platform. For attachments; until no reference to the attached file exists.

### Scalability & Reliability - Distributed Computing

#### Scalable Data Engine Worker Pool

Decube’s engine dynamically scales based on workload, provisioning additional compute workers under load while minimizing resource usage during idle periods. The distributed architecture also ensures fault tolerance against node-level failures.

#### High Performance & Reliable Job Scheduling

Our custom-built job scheduler, written in Rust, delivers:

* **At-most-once scheduling** to prevent duplicate task execution.
* **At least once execution** to maintain coverage for critical workloads.

Understand further how these models are handled refer to decube architecture.


# Deployment Methods

This document explains how Decube is architected to deliver scalability, security, and flexibility across varying customer requirements.

Decube supports three deployment models: Multi-Tenant SaaS, Single-Tenant SaaS, and Bring Your Own Cloud (BYOC). Regardless of the deployment type, customer data security and metadata integrity are always core priorities.

## Deployment Models - High-Level Comparison

| Feature                       | Multi-Tenant SaaS                 | Single-Tenant SaaS               | Bring Your Own Cloud (BYOC)                                    |
| ----------------------------- | --------------------------------- | -------------------------------- | -------------------------------------------------------------- |
| **Control Plane**             | Shared (Decube-hosted)            | Shared (Decube-hosted)           | Shared (Decube-hosted)                                         |
| **Data Plane**                | Shared                            | Dedicated per customer           | Deployed in customer’s cloud                                   |
| **Data Residency**            | Decube Cloud (AWS)                | Decube Cloud (AWS)               | Customer’s Cloud                                               |
| **Isolation Level**           | Logical isolation                 | Logical + infra isolation        | Full physical and logical isolation                            |
| **Ideal For**                 | Startups, SMBs, general use cases | Enterprises needing more control | Highly regulated industries, on-premise or data-locality needs |
| **Setup Complexity**          | Minimal                           | Minimal                          | High (requires cloud deployment setup)                         |
| **Customer-Controlled Infra** | No                                | No                               | Yes (data plane fully controlled)                              |

To see the details for each deployment method, please go to the respective pages:

* [SaaS](/security-and-infrastructure/deployment-methods/saas-multi-tenant) (Multi-Tenant)
* [SaaS](/security-and-infrastructure/deployment-methods/saas-multi-tenant) (Single-Tenant)
* [Bring-Your-Own-Cloud (BYOC)](/security-and-infrastructure/deployment-methods/bring-your-own-cloud-byoc)


# SaaS (Multi-Tenant)

In this model, both the control and data planes are hosted and managed by Decube, and shared across multiple customers.

### SaaS Deployment - Simplified Block Diagram

**Diagram: SaaS - Simplified Block Diagram**

<figure><img src="/files/TGJ6FKXbAx8WzraNzIuS" alt=""><figcaption></figcaption></figure>

In this model, both the control and data planes are hosted and managed by Decube, and shared across multiple customers. Customers connect their data sources securely to enable metadata collection and observability.

#### Components:

* **Control Plane**
  * Handles subscriptions, access control, email notifications, routing, licensing, and user management
* **Multi-Tenant Data Plane**
  * Metadata API
  * Metadata Storage
  * Distributed Job Scheduler
  * Metadata Collector
  * Data Quality Monitoring
* **Customer Data Source**
  * Remains in the customer’s environment (e.g., cloud data warehouses like Snowflake, BigQuery)

#### How it Works:

* Metadata is collected using secure, read-only connectors.
* While collecting metadata and operational metrics, Decube may process limited personal data solely for monitoring and cataloging purposes.
* All operations are performed securely, using credentials provided by the customer.

#### Data Security:

* Communication is fully encrypted via TLS.
* Credentials are stored with dual-layer encryption and are inaccessible to Decube staff.


# SaaS (Single-Tenant)

This model provides a dedicated data plane for each customer while retaining the shared control plane hosted by Decube.

### Single-Tenant SaaS - Simplified Block Diagram

**Diagram: Single Tenant SaaS - Simplified Block Diagram**

<figure><img src="/files/cCwR4Vn6fbFW4vvaaFCV" alt=""><figcaption></figcaption></figure>

This model provides a dedicated data plane for each customer while retaining the shared control plane hosted by Decube.

#### Components:

* **Control Plane** (same as SaaS)
* **Single-Tenant Data Plane** (dedicated to one customer)
  * Metadata API
  * Metadata Storage
  * Distributed Job Scheduler
  * Metadata Collector
  * Data Quality Monitoring

#### How it Works:

* A dedicated infrastructure ensures isolation and compliance adherence.
* Metadata is securely collected using the same read-only methods as in Multi-Tenant SaaS.
* Ideal for customers with stricter data handling policies or compliance requirements.

#### Data Security:

* Logical and infrastructure-level isolation is maintained.
* Data remains within the customer’s environment unless explicitly configured otherwise.


# Bring-Your-Own-Cloud (BYOC)

In the BYOC model, Decube’s data plane is fully deployed in the customer’s own cloud for complete control and compliance with data locality requirements.

### Bring Your Own Cloud- BYOC Simplified Block Diagram

**Diagram: Bring Your Own Cloud - BYOC Simplified Block Diagram**

<figure><img src="/files/tgkvaGbhQSGF9JumAWax" alt=""><figcaption></figcaption></figure>

In the BYOC model, Decube’s data plane is fully deployed in the customer’s own cloud for complete control and compliance with data locality requirements.

#### Components:

* **Control Plane** (Decube-hosted)
* **Customer Cloud Environment**
  * Data Plane (deployed in customer's infrastructure)
    * Metadata API
    * Metadata Storage
    * Distributed Job Scheduler
    * Metadata Collector
    * Data Quality Monitoring
* **Customer Data Source**
  * Remains within the customer’s environment.

#### How it Works:

* All processing (metadata collection, quality monitoring) happens within the customer's cloud.
* Only essential metadata (e.g., configurations, license info) is synced back to Decube’s control plane.

#### Data Security:

* No sensitive metadata or raw data leaves the customer’s environment.
* This model is ideal for organizations with strict data governance or residency requirements.

To learn more about how data is handled refer to data policy.


# Data Policy

Here's how we handle your data.

We practice a strict policy of not egressing any actual data from our customer's database, **only** metadata is scanned.

The only exceptions to this is:

* When you click **Accept** on the Profiler within an Asset Detail as some data will be egressed to run the profiler. This is also an **opt-in-only** feature.

#### How do we handle the data that we stored?

* Your metadata collected will only be stored up to 20 days.
* Data egressed for profiling purposes within Table Overview and Field Statistics will be deleted automatically at 12.00 AM UTC everyday.


# Snowflake

Adding Snowflake to your decube connections helps your team to find relevant datasets, understand their quality via incident monitoring and apply governance policies via our data catalog.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, View
* **Sync Tags** — syncs Snowflake tags to assets in the Catalog
* **Sync Objects Descriptions** — syncs object descriptions from Snowflake to the Catalog
* **Sync Column Keys & Constraints** — syncs column keys and constraints (NOT NULL, UNIQUE, PRIMARY KEY, FOREIGN KEY) from Snowflake datasets to the Catalog
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **External Table** — external tables, enabling queries on data in external storage (e.g., S3, ADLS) without loading it into the database
* **View Table** — view tables, which are virtual tables based on SQL queries

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **View Table Lineage** — tracks virtual tables (views) and their data dependencies
* **External Table Lineage** — tracks the relationship between raw files in cloud storage (S3/ADLS) and the virtualized relational schema in your data platform
* **SQL Query Lineage** — maps data movement through SQL queries (SELECT, JOIN, INSERT, etc.)
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Configurable Collection
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* Foreign Key Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

Below are the steps to Connect snowflake to Decube with key pair authentication method.

### Key Pair

Refer to the [Snowflake documentation](https://docs.snowflake.com/en/user-guide/key-pair-auth) for more information on how to generate a key pair. Please provide only the **unencrypted version** of private keys as the key is encrypted on Decube's end. The following credentials are required upon adding new connection:

* Username
* Public Key File
* [Account Identifier](#account-identifier)
* Warehouse Name
* Role Name

<figure><img src="/files/X6ESbgeBGrdfYfq7KEWQ" alt=""><figcaption><p>Snowflake</p></figcaption></figure>

The `source name` will be for you to differentiate and recognize particular sources within the decube application.

### Prerequisite

To ensure a smooth experience configuring the connection.

1. The user `decubeuser` and role `decuberole` is created and given the proper privileges for monitoring.
2. Steps 4 or 5 below needs to be repeated for every `database` that needs to be monitored.

### Account Identifier

If your Snowflake UI differ, please refer to [Snowflake Documentation](https://docs.snowflake.com/en/user-guide/admin-account-identifier#format-1-preferred-account-name-in-your-organization) on how to get your Account Identifier.

Account Identifier can be found by clicking on your profile icon in Snowflake and going to account details.

<figure><img src="/files/FbEiNWe4MtWTPaIiLlCB" alt=""><figcaption></figcaption></figure>

Copy the Account Identifier which is shown below. This should look like `<org-name>-<account-name>`

<figure><img src="/files/y1SMmuHpI4Sh5zXZ1MiB" alt=""><figcaption></figcaption></figure>

### Configuring User, Role and Privileges

1\. On a Snowflake worksheet, copy the commands below and modify as necessary. We have the user called `DECUBEUSER` and role called DECUBEROLE

```sql
set role_name = 'DECUBEROLE';
set user_name = 'DECUBEUSER';
set public_key = 'changethispublickey'; -- Change this to the public key you generated

-- !!Choose an existing warehouse name if you don't want to create a new warehouse!!
set warehouse_name = 'DECUBE_WH';

-- Creates the role and user and grant the role to the user
CREATE ROLE IF NOT EXISTS identifier($role_name);
CREATE USER IF NOT EXISTS identifier($user_name) DEFAULT_ROLE = $role_name;
ALTER USER identifier($user_name) SET RSA_PUBLIC_KEY = $public_key;
GRANT ROLE identifier($role_name) TO USER identifier($user_name);

-- Grants the base SNOWFLAKE database to the role
grant imported privileges on database "SNOWFLAKE" to role identifier($role_name);

-- This will create a warehouse if the chosen warehouse does not exists
CREATE warehouse IF NOT EXISTS identifier($warehouse_name)
warehouse_size = xsmall
warehouse_type = standard
auto_suspend = 5
auto_resume = true
initially_suspended = true
max_concurrency_level = 30
statement_timeout_in_seconds = 1800
statement_queued_timeout_in_seconds = 1200;

-- This grants the role access to the warehouse
grant USAGE on warehouse identifier($warehouse_name) to role identifier($role_name);
```

{% hint style="warning" %}
`statement_timeout_in_seconds` controls how long Snowflake allows a single query to run before terminating it. This affects all queries Decube runs against this warehouse, including metric collection for monitors and profiling jobs.

The recommended value is `1800` (30 minutes), which matches Decube's profiler job timeout. The minimum recommended value is `300` (5 minutes). Adjust within this range based on the complexity of your expected queries — wider tables and larger datasets require more time.

If this warehouse is shared with other workloads, a higher timeout means longer-running statements from those jobs will also be permitted to run before Snowflake terminates them. If your warehouse handles mixed workloads, consider creating a dedicated warehouse for Decube instead.
{% endhint %}

2\. The `source` type for the database has to be known. To get this information, from Snowflake dashboard click on *`Data`* -> *`Databases`.* On the left panel, a list of `Databases` can be seen along with `Source`.

3\. If `Source` if *local*, modify `database_name`, copy into a worksheet and run the commands.

```sql
set database_name = 'changethisdatabase'; -- The database to grant access to
set role_name = 'DECUBEROLE'; -- or differently if you modified it previously

-- Read-only access to database
grant USAGE on database identifier($database_name) to role identifier($role_name);
grant USAGE on all schemas in database identifier($database_name) to role identifier($role_name);
grant USAGE on future schemas in database identifier($database_name) to role identifier($role_name);
grant SELECT on all tables in database identifier($database_name) to role identifier($role_name);
grant SELECT on future tables in database identifier($database_name) to role identifier($role_name);
grant SELECT on all views in database identifier($database_name) to role identifier($role_name);
grant SELECT on future views in database identifier($database_name) to role identifier($role_name);

-- Only if external tables are to be ingested into Decube
grant SELECT on all external tables in database identifier($database_name) to role identifier($role_name);
grant SELECT on future external tables in database identifier($database_name) to role identifier($role_name);

```

4\. Step 3 needs to be repeated for every `database` that needs to be monitored.


# Redshift

Adding Redshift to your decube connections helps your team to find relevant datasets, understand their quality via incident monitoring and apply governance policies via our data catalog.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, View, Data Job, Data Run, Data Task
* **Sync Objects Descriptions** — syncs object descriptions from Redshift to the Catalog
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **Configurable Collection** — selective ingestion of schemas/workspaces in Data Source Management
* **View Table** — view tables, which are virtual tables based on SQL queries
* **Stored Procedure** — stored procedures (precompiled SQL; listed as Data Jobs in Metadata)

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
* Job Failure

**Lineage**

* **View Table Lineage** — tracks virtual tables (views) and their data dependencies
* **SQL Query Lineage** — maps data movement through SQL queries (SELECT, JOIN, INSERT, etc.)
* **Foreign Key Lineage** — tracks relationships between tables via primary and foreign keys
* **Stored Procedure Lineage** — tracks data flow through stored procedures as they execute
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* External Table

**Lineage**

* External Table Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

### Allowing Access

To allow our connector to access your Redshift instance, you will need to either:

1. Allow public access
2. Connect through a SSH Bastion

### Allowing Public Access

You can still limit who can connect to your Redshift instance through security-group inbound rules when you enable public access.

Go to Actions > Modify publicly accessible setting

<figure><img src="/files/cw6aBZfKgdyWoYDs1xAU" alt=""><figcaption><p>Ref 1.1</p></figcaption></figure>

Check `Turn on Publicly accessible` and select an Elastic IP address

<div align="center"><figure><img src="/files/mtvxVoOcBhSlwuPGO7V7" alt=""><figcaption><p>Ref 1.2</p></figcaption></figure></div>

Navigate to the `Properties` tab

<figure><img src="/files/YQyhpSPrEpDyW86DQXYk" alt=""><figcaption><p>Ref 1.3</p></figcaption></figure>

Scroll down to the `Network and security settings` and click through to your security group

<figure><img src="/files/xSy14bSiXsmWgDYLIA1T" alt=""><figcaption><p>Ref 1.4</p></figcaption></figure>

Navigate to the `Inbound rules` tab and click `Edit inbound rules`

<figure><img src="/files/rR5b88NN5EMo0UGnHhji" alt=""><figcaption><p>Ref 1.5</p></figcaption></figure>

Click `Add rule` and in `Type` choose Redshift and in the `Source` section, add all of Decube's collector IPs.

{% hint style="info" %}
You will need to post-fix the IP with /32 to limit it to only that IP. I.e. `xxx.xxx.xxx.xxx/32`
{% endhint %}

<figure><img src="/files/mlsOPXdG37usIC3i5yLU" alt=""><figcaption><p>Ref 1.5</p></figcaption></figure>

{% hint style="warning" %}
Be careful with modifying inbound rule policies. It can affect connectivity within your own VPC if you remove existing rules.
{% endhint %}

### SSH Bastion

You can also use a SSH Bastion if enabling public access is not an option. Setting up a Bastion host is out of the scope of this guide but you can refer to [SSH Tunneling](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ssh-tunneling) guide for more information.

Once you have setup a Bastion host, modify your Redshift security group inbound rule (refer to Ref 1.5) to allow source connection from your Bastion host's private IP address instead.

### Connection Details

Connecting to decube is as easy as providing us with credentials to your Redshift database. At a minimum, we require

* `username`
* `password`
* `host address`
* `host port`
* `database name`

<figure><img src="/files/mlAJtayTd8DTGMMypkPw" alt=""><figcaption><p>Amazon Redshift</p></figcaption></figure>

The `source name` will be for you to differentiate and recognize particular sources within the decube application.

We strongly encourage you to create a decube **read-only** user for this credential purpose, which you can follow [these steps](#custom-user-for-decube).

### **Security Concerns**

### Custom User for decube

We highly recommend that you create a Read-Only user for decube. We have prepared a script that you may run on your Redshift database to create this user.

1. Create a New User for decube

```
CREATE USER decube PASSWORD 'a_new_password';
```

2\. Add access to SYSLOG to build lineage and ingest Stored Procedures.

```
ALTER USER decube WITH SYSLOG ACCESS UNRESTRICTED;
```

3. Add access to information\_schema.

```
GRANT USAGE ON SCHEMA information_schema TO decube;
GRANT SELECT ON ALL TABLES IN SCHEMA information_schema TO decube;
```

4\. You may need to run this per schema that you have based on the default behavior of the schema.

```
GRANT USAGE ON SCHEMA <schema_name> TO decube;
GRANT SELECT ON ALL TABLES IN SCHEMA <schema_name> TO decube;
```


# Google Bigquery

Adding Google Big Query to your decube connections helps your team to find relevant datasets, understand their quality via incident monitoring and apply governance policies via our data catalog.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, View
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **Configurable Collection** — selective ingestion of schemas/workspaces in Data Source Management
* **View Table** — view tables, which are virtual tables based on SQL queries

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **View Table Lineage** — tracks virtual tables (views) and their data dependencies
* **SQL Query Lineage** — maps data movement through SQL queries (SELECT, JOIN, INSERT, etc.)
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* External Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* External Table Lineage
* Foreign Key Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to Big Query requires following the steps below to completed.

<figure><img src="/files/V4o5m09escosD344AX5K" alt=""><figcaption><p>Google Big Query</p></figcaption></figure>

### Prerequisite

To ensure a smooth experience configuring the connection.

1. Having [access](https://cloud.google.com/iam/docs/granting-changing-revoking-access#iam-view-access-console) to the Project that contains the intended BigQuery.
2. An account that can view and create [Service Accounts](https://cloud.google.com/iam/docs/understanding-service-accounts)

### Creating decube Data Observability Role

A custom role for the decube service account is required to ensure we have correct access to the resources we need.

1. From the dashboard, click the hamburger (three lines indicating a menu), the top left of the browser next to the Google Cloud logo, and Search for `IAM and Admin`.
2. Find `Roles` and at the top, below the search bar, click *`+ Create Role`.*
3. Fill in `Roles` as below

* `Title`: decube data observability
* `ID`: decube
* `Role launch Stage`: General Availability

1. Click on *`Add permission`,* a popup will be displayed. Then on *`Enter property name or value`* type each of the permissions below, press enter and click on the checkbox to assign that permission to the decube data observability role.

```
bigquery.datasets.get
bigquery.datasets.getIamPolicy
bigquery.jobs.create
bigquery.jobs.get
bigquery.jobs.list
bigquery.jobs.listAll
bigquery.readsessions.create
bigquery.readsessions.getData
bigquery.readsessions.update
bigquery.routines.get
bigquery.routines.list
bigquery.tables.get
bigquery.tables.getData
bigquery.tables.list
resourcemanager.projects.get
storage.buckets.get
storage.buckets.list
storage.objects.get
storage.objects.list
```

1. When you are done assigning permissions, click `Add` to exit the popup, on `Assigned Permission` table, review the permissions granted to match above requirements.
2. Click *`Create`*.

### Enabling Reporting Functionality

Our reporting module requires additional permission and Google Cloud Asset API to be enabled. We suggest creating a new role and attaching it to the existing service account above.

```
cloudasset.assets.analyzeIamPolicy
cloudasset.assets.searchAllIamPolicies
cloudasset.assets.searchAllResources
```

{% hint style="warning" %}
For full report generation especially with multiple GCP projects, the service account will need additional role on the Organization level. Without these additional roles, the generated report will only contain queries from the current project

```
bigquery.jobs.list
bigquery.jobs.listAll
```

{% endhint %}

1. Go to GCP API and Services and click Library.

<figure><img src="/files/ISM2giRyCdYtcWI9JvmM" alt=""><figcaption></figcaption></figure>

2. Search for “cloud asset api”.

<figure><img src="/files/tFv5UTVdyrUoh9ZW3Foj" alt=""><figcaption></figcaption></figure>

3. Click Enable

<figure><img src="/files/pvKUckUnO5Z2WZDgfqIu" alt=""><figcaption></figcaption></figure>

4\. Go to `IAM and admin` and `Roles` section click `Create Role.`

<figure><img src="/files/uCPnZVrAdoindXnJJeJh" alt=""><figcaption></figcaption></figure>

5\. Grant the role with these permissions and click `Create`

```
cloudasset.assets.analyzeIamPolicy
cloudasset.assets.searchAllIamPolicies
cloudasset.assets.searchAllResources
```

<figure><img src="/files/grHSTEO7RXwdyaHDTTJ0" alt=""><figcaption></figcaption></figure>

6. Grant the Role to the service account attached with the Bigquery source.

<figure><img src="/files/LdwFFraANSimQNi8DWMs" alt=""><figcaption></figcaption></figure>

### Getting a Service Account JSON Key File

For decube to connect with BigQuery, a JSON key file is required from a Service Account with the proper Roles assigned to it.

1. From the dashboard, click the hamburger (three lines indicating a menu), the top left of the browser next to the Google Cloud logo, and Search for *`API's and Services`.*
2. On that page, click on *`Credentials`*, at the top, click *`+ Create Credentials`* and then select *`Service account`*.
3. Fill in the *`Service account details form`,* we recommend

* `Service account name`: decube,

1. Click *`Create and Continue`.*
2. On *`Grant this service account access to project`*, click on *`Select a role`,* the on the \_`Filter` \_ search, type *`decube data observability`* created previously and click to assign this role to the service account. Click Continue.
3. Skip *`Grant users access to this service account`* form and click `Done`.
4. On the Service Account page you will see a Service Account named decube, click on it.
5. On the options below the Service Account name, click on `Keys`
6. Click *`Add Key`* and then select *`Create new key`* from the drop-down.
7. For the `Key Type`, choose `JSON` and click Create.
8. You will automatically download the JSON file and can refer to the location of the file on your browser, save it somewhere easily accessible like your Desktop.


# Databricks

Adding Databricks to your decube connections helps your team to find relevant datasets, understand their quality via incident monitoring and apply governance policies via our data catalog.

{% hint style="info" %}
The Databricks connection supports connecting to the Unity Catalog, or the legacy Hive metastore.
{% endhint %}

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, Data Job, Data Run, Data Task
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **View Table** — view tables, which are virtual tables based on SQL queries

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
* Job Failure

**Lineage**

* **View Table Lineage** — tracks virtual tables (views) and their data dependencies
* **Databricks Job Lineage** — tracks data movement and transformations through Databricks pipeline jobs
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Configurable Collection
* External Table

**Lineage**

* External Table Lineage
* SQL Query Lineage
* Foreign Key Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

## Adding a Databricks connection on decube.

From the Add New Connection page, find the Databricks connector logo to go to the connection page.

<figure><img src="/files/zppAlndJ2iPo8YdiUDq4" alt=""><figcaption></figcaption></figure>

## Required Information

1. Workspace URL - [jump to section](#getting-the-workspace-url)
2. Personal Access Token - [jump to section](#getting-a-personal-access-token)
3. SQL Warehouse HTTP Path - [jump to section](#getting-the-sql-warehouse-http-path)

### 1. Getting the Workspace URL

Go to your workspace and copy the url in the browser bar.

It should look like `<some_values>.cloud.databricks.com` . The entire url highlighted in the screenshot is the **Workspace URL** to be added to decube's form.

<figure><img src="/files/208Pkp3Apo3lBysRrFEZ" alt=""><figcaption></figcaption></figure>

### 2. Getting a Personal Access Token

The full documentation from Databricks can be found [here](https://docs.databricks.com/dev-tools/auth.html).

1. Navigate to your Users Settings page.

<figure><img src="/files/2oxl8XqGajSNcEjWKZBg" alt=""><figcaption></figcaption></figure>

2. Click on `Generate New Token` after navigating to the `Access Tokens` tab.

<figure><img src="/files/DTOnJvFcOGlMV1n4P2xs" alt=""><figcaption></figcaption></figure>

3. Give your token a name and specify the lifetime of the token. We suggest not specifying a lifetime to ensure uninterrupted service.
4. Once you click `Generate` ensure that you note down the token somewhere as you cannot retrieve it again. This is the **Access Token** for decube's form.

### 3. Getting the SQL Warehouse HTTP Path

1. Go to your `SQL` Workspace as shown and then Navigate to the `SQL Warehouses` section.

<figure><img src="/files/BXdTsFAaFC6qQF0BXWVN" alt=""><figcaption></figcaption></figure>

2. Either create a new SQL warehouse (recommended) or choose an existing SQL Warehouse to be used with decube and navigate to the `Connection Details` tab. You will see both the Server hostname which should match your Workspace URL and the HTTP Path of the SQL Warehouse

{% hint style="warning" %}
We recommend creating a serverless Databricks SQL Warehouse. Other solutions may cause failure during metadata ingestion or data monitoring due to warehouse warm-up time.
{% endhint %}

<figure><img src="/files/IZBvTruO7rjj0ldIr4Zg" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
If there is any issues or questions related with connecting your data sources, please reach out to us via our Live chat for support.
{% endhint %}

### FAQ

#### If I have multiple catalogs in my Databricks, how will that be handled?

{% hint style="info" %}
The Catalog is the first layer of the object hierarchy, used to organize your data assets in the Unity Catalog. Read more about the Unity Catalog [here](https://docs.databricks.com/en/data-governance/unity-catalog/index.html).
{% endhint %}

Decube ingests all Databricks catalogs upon adding the connection into a single data source.


# Azure Synapse

Adding Synapse to your decube connections helps your team to find relevant datasets, understand their quality via incident monitoring and apply governance policies via our data catalog.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **Configurable Collection** — selective ingestion of schemas/workspaces in Data Source Management
* **External Table** — external tables, enabling queries on data in external storage (e.g., S3, ADLS) without loading it into the database

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **External Table Lineage** — tracks the relationship between raw files in cloud storage (S3/ADLS) and the virtualized relational schema in your data platform
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* View Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* View Table Lineage
* SQL Query Lineage
* Foreign Key Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to Azure Synapse requires credentials that can be easily found through this guide. These credentials include:

* `username`
* `password`
* `host`
* `database`

<figure><img src="/files/98AGsnLalOk6wwgeTCNr" alt=""><figcaption><p>Azure Data Lake Storage</p></figcaption></figure>

## Prerequisite

* Ensure the Azure account you logged into has access to Azure Synapse credentials.\\

## Getting the Credentials.

Log into your Azure account that has access to manage Synapse and click on the Synapse workspace of your choice. These are the information that you need to be aware of.

* **Dedicated SQL server URL (1)** or **Serverless SQL endpoint (2)**. Depending on where the database you want to connect with decube resides, we will need to input this information for Host.

These are highlighted and numbered on the relevant image below.

<figure><img src="/files/r7EqIXDPX1BZ74sJgZ3j" alt=""><figcaption><p>Retrieve the Dedicated SQL Server URL or Serverless SQL endpoint from this screen.</p></figcaption></figure>

Click on the Synapse workspace of your choice. These are the credentials that we will need from here.

1. Username
2. Password
3. Database name

<figure><img src="/files/B6iRl3srjRlJMHWURLHN" alt=""><figcaption></figcaption></figure>

1. Select the master database on the `Use database` drop-down and create this login, we recommend calling it `decube_login` .

Save this login value and password value, the login will be for the `Username` and `Password` credentials respectively.

```sql
CREATE LOGIN decube_login WITH PASSWORD = 'strong_password_here!';
GO
```

2. Select the database you wish to connect with decube on with the `Connect to` and `Use database` drop-down. Create a user with the mapping for the login. We recommend naming it `decube_user` .

Save the database name as we will need that as the `Database` credential.

```sql
CREATE USER decube_user
FOR LOGIN decube_login;
GO
```

3. (a) Grant the user permission to read the data based on your Synapse SQL pool type.

```sql
ALTER ROLE db_datareader ADD MEMBER decube_user;
GO
```

(b) If the database is in a Dedicated SQL pool:

```sql
EXEC sp_addrolemember 'db_datareader', 'decube_user';
GO
```

4. Add the permissions below to use Data Quality and Profiler.

```sql
GRANT VIEW DATABASE STATE TO decube_user;
GRANT VIEW DEFINITION TO decube_user;
GO
```

5. Grant Storage Credential Access **(Serverless SQL Pools Only)**

Grant the user permission to authenticate against Azure Storage.

Ensure the tables are created with the items below:

* A `Database Scoped Credential` created in the database (using Managed Identity).
* An `External Data Source` attached to the Database Scoped Credential.
* The `External Tables` have been successfully created using the External Data Source.

Run the command below, replacing `[Your_Credential_Name]` with the actual credential of your `External Data Source`:

```sql
GRANT REFERENCES ON DATABASE SCOPED CREDENTIAL::[Your_Credential_Name] TO decube_user;
GO
```

> For more information, refer to Microsoft's official documentation on [using external tables](https://learn.microsoft.com/en-us/azure/synapse-analytics/sql/develop-tables-external-tables) and [accessing external tables](https://learn.microsoft.com/en-us/azure/synapse-analytics/sql/develop-storage-files-overview) with Serverless Synapse.


# Microsoft Fabric (BETA)

Adding Microsoft Fabric to your decube connections helps your team find relevant datasets and apply governance policies via the data catalog

Microsoft Fabric integrates with decube through a service principal, giving your team metadata visibility and governance coverage over Fabric workspaces directly in the data catalog.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* External Table
* Stored Procedure
  {% endtab %}
  {% endtabs %}

{% hint style="warning" %}
The connector may be rate-limited when scanning more than 1,000 workspaces.
{% endhint %}

## Connection Requirements

Connecting Fabric requires credentials you generate in Microsoft Entra ID:

* `Tenant ID`
* `Client ID`
* `Client Secret`

<figure><img src="/files/yj9WnmcGq2hAGhFYuEdq" alt=""><figcaption></figcaption></figure>

### Prerequisites

1. Access to Microsoft Entra ID to create service principals.
2. A Microsoft Fabric-enabled workspace.
3. Admin access (Fabric Administrator role) to change tenant settings.

### Creating a service principal and security group

1. Go to **Microsoft Entra ID**. Under the **Manage** tab on the left, select **App Registrations**.
2. Click **New Registration** at the top of the page.

<figure><img src="/files/y2Qd3c7Otp37LfukttbE" alt=""><figcaption></figcaption></figure>

3. Enter `decube` as the name, then click **Register**.
4. On the **Overview** tab, note the `Client ID` and `Tenant ID` — you'll need these when configuring the connector in decube.

<figure><img src="/files/1g5UY1An9O4pFhJSoj3W" alt=""><figcaption></figcaption></figure>

5. On the left tab, click **Certificates & secrets**, then **New client secret**. Set the description to `decube client secret` and the expiry date to align with your contract term.

<figure><img src="/files/DLocKnd7TMOEG7C9cxKF" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
Copy the client secret value immediately and store it somewhere secure, such as Azure Key Vault. Microsoft Entra ID doesn't display it again, and you'll need it to register the connector in decube.
{% endhint %}

6. Return to the Microsoft Entra ID main page and click **Groups** on the left, then **New Group**.

<figure><img src="/files/jtD5iSR7J3sEg1rHTCMS" alt=""><figcaption></figcaption></figure>

7. Set the group type to **Security**.
8. Click **Members**, then search for the service principal you created (`decube`).
9. Click **Select** to add the service principal to the group.
10. Click **Create**. You now have the Tenant ID, Client ID, and Client Secret needed for the integration.

### Granting the service principal access to Microsoft Fabric

1. On the Microsoft Fabric workspace page, click the **Settings** gear icon in the top-right corner, then select **Admin Portal** under **Governance and Insights**.
2. In the Admin Portal, select **Tenant Settings** in the left sidebar.
3. Search for **Service principals can call Fabric public APIs**.
4. Enable the setting and select **Specific security groups**, then add the group you created earlier.

<figure><img src="/files/LBxAMd0jnoDH9wd16G3r" alt=""><figcaption></figcaption></figure>

5. Enable the following settings as well, under **Admin API**:

* **Service principals can access read-only admin APIs**
* **Enhance admin APIs responses with detailed metadata**
* **Enhance admin APIs responses with DAX and mashup expressions**

### Granting Workspace Access to the Service Principal

Please follow these steps for all of the Workspaces that you want ingested. A user with `Contributor` permissions on the workspace is required to perform this step.

1. In the Microsoft Fabric portal, select `Workspaces` from the left sidebar to view available workspaces.
2. Hover over the workspace name that you want ingested and select the `...` ellipsis menu that appears. Click `Workspace access`. (Alternatively, if you are already inside the workspace, you can select `Manage access` in the top right corner).

<figure><img src="/files/JRTV6afiBDs4aZ0jvZE9" alt=""><figcaption></figcaption></figure>

3. The `Manage access` panel will open on the right side of the screen. Click `Add people or groups`. In the search field, search for the name of the service principal created earlier (in this case, `decube`).
4. Change the access role drop down to `Contributor`, then finally, click on `Add`.

<figure><img src="/files/HMIWkiXmylCPuLoJopnX" alt=""><figcaption></figcaption></figure>


# Clickhouse

Adding Clickhouse to your decube connections helps your team find relevant datasets and apply governance policies via the data catalog

{% hint style="info" %}
This connector will be released to BETA soon. Stay tuned.
{% endhint %}


# PostgreSQL

Add PostgreSQL as a decube connection and help your team discover, document and monitor their data assets to drive data-driven insights and decisions.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, View
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **Configurable Collection** — selective ingestion of schemas/workspaces in Data Source Management
* **View Table** — view tables, which are virtual tables based on SQL queries

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **View Table Lineage** — tracks virtual tables (views) and their data dependencies
* **Foreign Key Lineage** — tracks relationships between tables via primary and foreign keys
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* External Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* External Table Lineage
* SQL Query Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to decube is as easy as providing us with credentials to your PostgreSQL database. At a minimum, we require:

* `username`
* `password`
* `host address`
* `host port`
* `database name`

<figure><img src="/files/krmj4GIVvyeKBtkWnyYN" alt=""><figcaption><p>PostgreSQL</p></figcaption></figure>

The `source name` will be for you to differentiate and recognize particular sources within the decube application.

We strongly encourage you to create a decube **read-only user** for this credential purpose, which you can follow [here](#custom-user-for-decube).

### **Security Concerns**

If access to your database is protected by security measures, we allow for connecting via [SSH Tunneling](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ssh-tunneling) or you could [whitelist our IP](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ip-whitelisting). See more [here.](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access)

### Custom User for decube

We highly recommend that you create a **Read-Only** user for decube. We have prepared a script that you may run on your PostgreSQL database to create this user.

1. Create a New User for decube

```
CREATE ROLE decube WITH LOGIN ENCRYPTED PASSWORD 'a_new_password';
```

2\. Allow connection to the database

```
GRANT CONNECT ON DATABASE "your_db_name" TO decube;
```

3\. We need access to information\_schema

```
GRANT USAGE ON SCHEMA information_schema TO decube;
GRANT SELECT ON ALL TABLES IN SCHEMA information_schema TO decube;
```

4\. You may need to run this per schema that you have based on the default behavior of the schema.

```
GRANT USAGE ON SCHEMA <schema_name> TO decube; 
GRANT SELECT ON ALL TABLES IN SCHEMA TO decube;
```

If you want all of the tables in all your schemas to be included in decube, run the following SQL script:

```sql
DO $do$
DECLARE
    sch text;
BEGIN
    FOR sch IN SELECT nspname FROM pg_namespace where nspname !~* 'pg|information_schema'
    LOOP
        EXECUTE format($$ GRANT USAGE ON SCHEMA %I TO decube $$, sch);
    EXECUTE format($$ GRANT SELECT ON ALL TABLES IN SCHEMA %I TO decube $$, sch);
    END LOOP;
END;
$do$;
```


# MySQL

Add MySQL as a decube connection and help your team discover, document and monitor their data assets to drive data-driven insights and decisions.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, View
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **View Table** — view tables, which are virtual tables based on SQL queries

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **View Table Lineage** — tracks virtual tables (views) and their data dependencies
* **Foreign Key Lineage** — tracks relationships between tables via primary and foreign keys
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Configurable Collection
* External Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* External Table Lineage
* SQL Query Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to decube is as easy as providing us with credentials to your MySQL database. At a minimum, we require

* `username`
* `password`
* `host address`
* `host port`
* `database name`

<figure><img src="/files/PiFGOkFDPZZ5vsQw8Yor" alt=""><figcaption><p>MySQL</p></figcaption></figure>

The `source name` will be for you to differentiate and recognize particular sources within the decube application.

We strongly encourage you to create a decube **read-only** user for this credential purpose, which you can follow [here](#custom-user-for-decube).

## **Security Concerns**

If access to your database is protected by security measures, we allow for connecting via [SSH Tunneling](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ssh-tunneling) or you could [whitelist our IP](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ip-whitelisting). See more [here.](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access)

### Configuring a Custom MySQL User

Creating a dedicated database user for Decube is the recommended approach. This follows the principle of least privilege, ensuring Decube has only the specific permissions it needs to monitor your data without risking broader database security.

#### 1. Create the Decube User

Run the following command to initialize the user. Replace `<host>` with Decube's IP (or `%` for any host) and set a strong password.

```sql
CREATE USER 'decube'@'<host>' IDENTIFIED BY '<password>';
```

#### 2. Assign Necessary Permissions

Decube requires specific privileges to read metadata and monitor table health.

**Data Access (SELECT)**

You must grant `SELECT` access to the tables you wish to observe. Use a wildcard (`*`) for the entire database or specify individual tables.

* For the entire database:

  ```sql
  GRANT SELECT ON <database>.* TO 'decube'@'<host>';
  ```
* For specific tables only:

  ```sql
  GRANT SELECT ON <database>.<table> TO 'decube'@'<host>';
  ```

**View Metadata Access (SHOW VIEW)**

To monitor and process MySQL Views, Decube requires the `SHOW VIEW` privilege at the database level. Without this, Decube cannot ingest view definitions.

```sql
GRANT SHOW VIEW ON <database>.* TO 'decube'@'<host>';
```

## Use your own SSL CA Cert

<figure><img src="/files/761uZEH37X7tiW6I7vRd" alt=""><figcaption><p>Upload SSL CA File</p></figcaption></figure>

If your database enforces SSL connection, you can provide your own CA Cert file by choosing `Use Customer CA Cert`and uploading a CA Cert.

This works for AWS RDS Aurora MySQL instances as well. To see where to retrieve your Aurora certificates, please use this [guide from AWS](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/UsingWithRDS.SSL.html)


# SingleStore

Add SingleStore as a decube connection and help your team discover, document and monitor their data assets to drive data-driven insights and decisions.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **Foreign Key Lineage** — tracks relationships between tables via primary and foreign keys
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* View Table Lineage
* External Table Lineage
* SQL Query Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to decube is as easy as providing us with credentials to your SingleStore database. At a minimum, we require

* `Username`
* `Password`
* `Host Address`
* `Host Port`
* `Database`

<figure><img src="/files/xwah9o7YU3g5fghhxANW" alt=""><figcaption><p>SingleStore</p></figcaption></figure>

The `source name` will be for you to differentiate and recognize particular sources within the decube application.

We strongly encourage you to create a decube **read-only** user for this credential purpose, which you can follow [here](#custom-user-for-decube).

### **Security Concerns**

You may need to configure your SingleStore database to [whitelist our IP](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ip-whitelisting). You can do so under the Workspace > Firewall

### **Custom User for decube**

A custom user would allow for a granular configuration of the user on your database and your connection to decube.

1. Create the user, and change `host` and `password` accordingly.

```bash
CREATE USER 'decube' IDENTIFIED BY '<password>';
```

1. Grant access to the decube user.

```bash
GRANT SELECT ON <database>.* TO 'decube'@'%';
```


# Microsoft SQL Server

Add SQL Server as a decube connection and help your team discover, document and monitor their data assets to drive data-driven insights and decisions.

{% hint style="info" %}
You can also use this document to connect **Azure SQL Server** to Decube.
{% endhint %}

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **Foreign Key Lineage** — tracks relationships between tables via primary and foreign keys
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* View Table Lineage
* External Table Lineage
* SQL Query Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to decube is as easy as providing us with credentials to your Microsoft SQL database. At a minimum, we require

* `username`
  * If connecting to AzureSQL using SSH, use `username@<servername>.database.windows.net`
* `password`
* `host address`
  * If connecting to AzureSQL, use fully qualified name e.g. `<servername>.database.windows.net`
* `host port`
* `database name`

<figure><img src="/files/E7h7ep7tqkjwhDJdjvEz" alt=""><figcaption><p>Microsoft SQL Server</p></figcaption></figure>

The `source name` will be for you to differentiate and recognize particular sources within the decube application.

We strongly encourage you to create a decube **read-only** user for this credential purpose, which you can follow [here](https://github.com/DecubeIO/decube_docs/blob/master/databases/mssql.md#custom-user-for-decube).

### **Security Concerns**

If access to your database is protected by security measures, we allow for connecting via [SSH Tunneling](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ssh-tunneling) or you could [whitelist our IP](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access/ip-whitelisting). See more [here.](https://docs.decube.io/how-to-connect-data-sources/enabling-vpc-access)

{% hint style="info" %}
To connect to Azure SQL through an SSH tunnel, you must explicitly provide the server context. Azure's gateway requires the following configuration to route the request correctly:

* Host: Use the fully qualified domain name: `<servername>.database.windows.net`.
* Username: Use the format `username@<servername>.database.windows.net`.
  {% endhint %}

### **Custom User for decube**

A custom user would allow for a granular configuration of the user on your database and your connection to decube.

1. Create a New User for decube

```sql
CREATE LOGIN [decube] WITH PASSWORD=N'your_password';
```

2. Execute the following SQL statement to create a new database user and map it to the login created in the previous step

```sql
USE your_db_name;
CREATE USER [decube] FOR LOGIN [decube];
```

3. Add the user to role db\_datareader

```sql
EXEC sp_addrolemember N'db_datareader', N'decube'
```


# Oracle

Add Oracle as a decube connection and help your team discover, document and monitor data assets to drive data-driven insights and decisions.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column
* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **Configurable Collection** — selective ingestion of schemas/workspaces in Data Source Management
* **View Table** — view tables, which are virtual tables based on SQL queries

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift

**Lineage**

* **View Table Lineage** — tracks virtual tables (views) and their data dependencies
* **Foreign Key Lineage** — tracks relationships between tables via primary and foreign keys
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* External Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure

**Lineage**

* External Table Lineage
* SQL Query Lineage
* Stored Procedure Lineage
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to decube is as easy as providing us with read-only credentials to your Oracle database. At a minimum, we require

* `Username`
* `Password`
* `Host`
* `Port`
* `Database`

<figure><img src="/files/AG4xCldIolqKzN9C6snL" alt=""><figcaption><p>Oracle</p></figcaption></figure>

### Create User for Decube

A custom user would allow for a granular configuration of the user on your database and your connection to decube.

1. Create the user, and change password accordingly.

```sql
CREATE USER 'decube' IDENTIFIED BY '<password>';
```

2. Grant `read-only` privileges on all data dictionary views

```sql
GRANT SELECT_CATALOG_ROLE TO 'decube';
```

3. Grant access of tables and views to the decube user. Here `table` or `view` can be `*` if you're using decube to observe the whole database. For only a few tables, you'll have to run this command multiple times changing the `table` or `view` value.

```sql
GRANT CONNECT TO decube;
GRANT SELECT ON <schema>.<table> TO decube;
GRANT SELECT ON <schema>.<view> TO decube;
```


# SAP HANA (BETA)

Add SAP HANA as a decube connection and help your team discover, document and monitor their data assets to drive data-driven insights and decisions.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, View
* **View Table** — view tables, which are virtual tables based on SQL queries
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* External Table
* Stored Procedure
* Calculation View
* Analytic View
* Attribute View
  {% endtab %}
  {% endtabs %}

## Connection Requirements

Connecting to decube requires credentials to your SAP HANA instance. At a minimum, we require:

* `username`
* `password`
* `host address`
* `host port`
* `database / tenant name`

The `source name` will be for you to differentiate and recognize particular sources within the decube application.

We strongly encourage you to create a decube **read-only user** for this credential purpose. Follow the steps below to set up a custom user with least-privilege access.

### Custom User for decube

Creating a dedicated database user for decube follows the principle of least privilege, giving decube only the permissions it needs to monitor your data without risking broader database security.

1. Create the decube user

Log in to your SAP HANA instance with an administrative account (e.g., `SYSTEM` or a user with `USER ADMIN` privileges) and run:

```sql
-- Create the dedicated service user, keep note of the username and password
CREATE USER DECUBE PASSWORD "<password>" NO FORCE_FIRST_PASSWORD_CHANGE;

-- Disable password expiration for uninterrupted ingestion
ALTER USER DECUBE DISABLE PASSWORD LIFETIME;
```

2\. Grant the required permissions

decube requires `SELECT METADATA` granted to the `DECUBE` user, for every schema you want to ingest:

```sql
-- Grant metadata read access for each schema you wish to ingest.
GRANT SELECT METADATA ON SCHEMA "<schema_name>" TO DECUBE;
```

See the SAP documentation on [SELECT METADATA](https://help.sap.com/docs/SAP_HANA_PLATFORM/4fe29514fd584807ac9f2a04f6754767/20f674e1751910148a8b990d33efbdc5.html) for more detail.

## Encryption & SSL Configuration

decube supports encrypted SSL connections to SAP HANA instances via the **Enable SSL** toggle.

If your enterprise database or cloud environment enforces custom CA verification, select **Use Customer CA Cert** and upload your `.pem` or `.crt` certificate file.


# dbt (Cloud Version)

Adding dbt to your decube connections helps you monitor your transformations via our Data Quality model and see metadata on your dbt models and jobs directly synced to the Data Catalog.

{% hint style="info" %}
This documentation is for the cloud version of dbt which is dbt Cloud. To see documentation on how to connect the open source version of dbt, please check out the documentation for dbt core [here](/transformation-tools/dbt-core).
{% endhint %}

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Data Job, Data Run, Data Task

**Data Quality Monitors**

* Job Failure
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
  {% endtab %}
  {% endtabs %}

dbt Cloud can map lineage relationships to upstream and downstream objects from the following connectors:

* Upstream Connectors: postgresql, redshift, snowflake, bigquery, mysql
* Downstream Connectors: postgresql, redshift, snowflake, bigquery, mysql

## Connection Requirements

### Minimum Requirements

{% hint style="success" %}
This connection requires [Service Account Tokens](https://docs.getdbt.com/docs/dbt-cloud-apis/service-tokens) to be generated in dbt Cloud, which is only accessible to dbt accounts on *Team* and *Enterprise* plans.
{% endhint %}

* Service Account API Key
* Account ID
* Tenancy Type
* Access URL (if non DBT multi-tenant)
* Discovery URL (if non DBT multi-tenant)

### Adding a dbt connection on decube.

* From the `My Account` page on decube platform, select on the dbt tile to be brought to the dbt connection form.

<figure><img src="/files/G0HwNAtqKm8vqu9w0NhF" alt=""><figcaption><p>DBT</p></figcaption></figure>

1. Account ID

* To get your Account ID, you can refer to the numbers at the end of the `accounts` path component of the URL (`https://cloud.getdbt.com/settings/accounts/{account_id}`).

2. API Key

* To get your API Key, please read on to the next section "`Retrieve Service Token from dbt account settings`" to generate your service token for your dbt account.

3. DBT Tenancy Type

* Usually this will be US Multi Tenant for most customers. To verify, check the domain you log in to your DBT with
  * US Multi Tenant - cloud.getdbt.com
  * EMEA Multi Tenant - emea.dbt.com
  * APAC Multi Tenant - au.dbt.com

If your DBT domain contains other URLs, you must choose `Other` as the tenancy type and provide item 4 and 5 mentioned below:

4. Access URL (For `Other` tenancy type)

* Navigate to your DBT Account `Settings` and below the Account Information pane, you can find the Access URL. You can follow the guide from DBT [here](https://docs.getdbt.com/docs/cloud/about-cloud/access-regions-ip-addresses#api-access-urls)

<figure><img src="/files/cClveryR8XHIZWxf0kLD" alt=""><figcaption></figcaption></figure>

5. Discovery URL (For `Other` tenancy type)

* Navigate to your DBT Account `Settings` and below the Account Information pane, you can find the Discovery API URL. You can follow the guide from DBT [here](https://docs.getdbt.com/docs/cloud/about-cloud/access-regions-ip-addresses#api-access-urls)

<figure><img src="/files/zkfWvevbPXbKVAyIK4a0" alt=""><figcaption></figcaption></figure>

### Retrieve Service Token from dbt account settings

1. Log into your dbt account. From the `Dashboard`, click on `gear icon` on the top right corner as shown in image, to navigate to Account Settings.

<figure><img src="/files/CQ2T0g4oZgkf7BJnAHBu" alt=""><figcaption><p>Click on the Account Settings on the top right of the dbt Dashboard.</p></figcaption></figure>

2. Within Account Settings, find the setting for `Service Tokens` on the left sidebar.

<figure><img src="/files/Mma19sgHNH1VE0BDnCvQ" alt=""><figcaption><p>Find Service Tokens in the Account Settings.</p></figcaption></figure>

3. Within the Service Tokens section, click on `+ New Token`.

<figure><img src="/files/zdyEGeqMnoQ0STFFb6SO" alt=""><figcaption><p>Add a new token under the Service Tokens section.</p></figcaption></figure>

4. Within the `New Service Token` page, you can set any token name, such as "decube". Then, click on the `+ Add` button.

<figure><img src="/files/SvdSLcPd15rtC44tH3gV" alt=""><figcaption><p>Give a name and add the permissions.</p></figcaption></figure>

5. Select `Job Admin` and add all the Projects that you would like to connect decube with. Then, click on `Save` at the bottom right.

<figure><img src="/files/OX9MY4cDGUOdhM2CY6cP" alt=""><figcaption><p>Setting the correct permissions for the Service Token.</p></figcaption></figure>

6. A token will be generated. Use the `Copy` button to copy the token into your clipboard.

<figure><img src="/files/eMu5y0m4xIRoC9IfLwf0" alt=""><figcaption><p>Copy the Service Token.</p></figcaption></figure>

7. Insert the Service token that you've copied into the "API Key" of the connection form, then test the connection. If it is successful, you can now add the name and submit the connection.

{% hint style="warning" %}
If there is any issues or questions related with connecting your data sources, please reach out to us via our Live chat for support.
{% endhint %}


# dbt Core

Connect your decube platform to dbt Core to see all data jobs in the Catalog and see end-to-end lineage.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Data Job, Data Run, Data Task

**Data Quality Monitors**

* Job Failure
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
  {% endtab %}
  {% endtabs %}

dbt Core can map lineage relationships to upstream and downstream objects from the following connectors:

* Upstream Connectors: postgresql, redshift, snowflake, bigquery, mysql
* Downstream Connectors: postgresql, redshift, snowflake, bigquery, mysql

## Connection Requirements

This documentation is on how to add a data source connection to dbt Core, which is the open source framework for dbt. If you are interested to connect to your dbt Cloud instance instead, please check out [this documentation](/transformation-tools/dbt) for dbt Cloud version.

{% hint style="info" %}
Important: Our system does not parse or collect metadata from past/old DBT runs.

In order for metadata to be collected and properly ingested into our system, the DBT data job must be re-run to get the same day data.
{% endhint %}

Integrating DBT Core with Decube involves reading files from an AWS S3 bucket, which shares similarities with how AWS S3 itself connects to the platform.

1. Set up an S3 bucket following the same procedure outlined in our documentation for [AWS S3](/datalake/s3).
2. Define folder partitions (details will be provided in the following section).
3. Upload the necessary files to those partitions.

A summary of steps to set up dbt core:

1. Set up an S3 bucket following the same procedure outlined in our documentation for [AWS S3](/datalake/s3).
2. Define folder partitions (details will be provided in the following section).
3. Upload the necessary files to those partitions.

Following these steps, the metadata collector will connect to the S3 bucket and retrieve the data.

## Minimum Requirement

{% hint style="info" %}
Currently, only S3 storage is supported for DBT Core under the "Storage Type" dropdown.
{% endhint %}

To connect your AWS Glue to decube, we will need the following information:

Choose authentication method:

a. [**AWS Identity**](#a.-aws-roles):

* Select AWS Identity
* Customer AWS Role ARN
* Path
* Region
* Storage Type
* Data source name

<figure><img src="/files/r13fuamEZISPlNxhPhHd" alt=""><figcaption><p>Connecting DBT Core using AWS Identity</p></figcaption></figure>

b. **AWS Access Key**:

* Access Key ID
* Secret Access Key
* Path
* Region
* Storage Type
* Data source name

{% hint style="info" %}
where 'Path' follows these format:\
s3://some-bucket\
s3://some-bucket/path-to-dbt-core

The path spec above will be created during the [Upload Project Files](#upload-project-files) step.
{% endhint %}

<figure><img src="/files/yis7r9nVETsNJ6zszTAi" alt=""><figcaption><p>Connecting DBT Core using AWS Access Key</p></figcaption></figure>

## Connection Options:

### a. AWS Roles

{% hint style="info" %}
This section will create a Customer AWS Role within your AWS account that has the right set of permission to access your data sources.
{% endhint %}

* Step 1: Go to your AWS Account → IAM Module → Roles
* Step 2: Click on **Create** **role**.

<figure><img src="/files/oqh7ru346tgegU5ag3yO" alt=""><figcaption></figcaption></figure>

* Step 3: Choose **Custom** **trust** **policy**.

<figure><img src="/files/69oNLlXEpHGoPBgwyZUV" alt=""><figcaption></figcaption></figure>

* Step 4: Specify the following as the trust policy, replacing `DECUBE-AWS-IDENTITY-ARN` and `EXTERNAL-ID` with values from [AWS Identities](/security-and-connectivity/aws-identities#generating-a-decube-aws-identity).

```
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "AWS": "<DECUBE-AWS-IDENTITY-ARN>"
            },
            "Action": "sts:AssumeRole",
            "Condition": {
                "StringEquals": {
                    "sts:ExternalId": "<EXTERNAL-ID>"
                }
            }
        }
    ]
}
```

* Step 5: Click next to proceed to attach policy.
* Step 6: Click on **Attach Policies** and Create Policy and choose **JSON Editor**. Input the following policy and press next, input the **policy name** of your choice and press **Create Policy**.

```
{
	"Version": "2012-10-17",
	"Statement": [
		{
			"Sid": "VisualEditor0",
			"Effect": "Allow",
			"Action": [
				"s3:GetObject",
				"s3:ListBucket",
				"s3:ListAllMyBuckets"
			],
			"Resource": [
				"arn:aws:s3:::{bucket-name}",
				"arn:aws:s3:::{bucket-name}/*"
			]
		}
	]
}
```

### b. Retrieving Access Keys from AWS

* Step 1: Login to AWS Console and proceed to IAM > User > Create User

<figure><img src="/files/nq4YDBjsM6Kr8QRs4vqN" alt=""><figcaption></figcaption></figure>

* **Extra Step:** Click on Attach Policies and Create Policy and choose JSON Editor input the following policy and press next, input the policy name of your choice and press Create Policy

```json
{
	"Version": "2012-10-17",
	"Statement": [
		{
			"Sid": "VisualEditor0",
			"Effect": "Allow",
			"Action": [
				"s3:GetObject",
				"s3:ListBucket",
				"s3:ListAllMyBuckets"
			],
			"Resource": [
				"arn:aws:s3:::{bucket-name}",
				"arn:aws:s3:::{bucket-name}/*"
			]
		}
	]
}
```

* Step 2: Search for the policy you created just now, select it and press **Next**.

<figure><img src="/files/lTQlFXHIGRTPEeAWx3kU" alt=""><figcaption></figcaption></figure>

* Step 3: Review and **Create user**.

<figure><img src="/files/bvvxnNxqR9khRuHvgrfs" alt=""><figcaption></figcaption></figure>

* Step 4: Navigate to the newly created user and click on `Create access key`

<figure><img src="/files/ECZ9A3zMvINAepOP5Zgj" alt=""><figcaption></figcaption></figure>

* Step 5: Choose `Application running outside AWS`

<figure><img src="/files/Kv9SUlwPFgKIDJC0iB6a" alt=""><figcaption></figcaption></figure>

* Step 6: Save the provided access key and secret access key. You will not be able to retrieve these keys again.

<figure><img src="/files/YIVfKpbkXJ6ZE8Wz0acN" alt=""><figcaption></figcaption></figure>

### AWS KMS

If the bucket intended to be connected to Decube is encrypted using a customer managed KMS key, you will need to add the AWS IAM user created above to the key policy statement.

1. Login to AWS Console and proceed to AWS KMS > Customer-managed keys.
2. Find the key that was used to encrypt the AWS S3 bucket.
3. On the Key policy tab, click on `Edit`

<figure><img src="/files/xFbswZjG4fJ8PzhQS723" alt=""><figcaption></figcaption></figure>

4. Assuming the user created is `decube-s3-datalake`

a. If there is not an existing policy attached to the key

```
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "Allow decube to use key",
            "Effect": "Allow",
            "Principal": {
                "AWS": [
                    "arn:aws:iam::<AWSAccountID>:user/{decube-s3-datalake}"
                ]
            },
            "Action": "kms:Decrypt",
            "Resource": "*"
        }
    ]
}
```

b. If there is an existing policy, append this section to the `Statement` array:

```
{
    "Statement": [
        {
            "Sid": "Allow decube to use key",
            "Effect": "Allow",
            "Principal": {
                "AWS": [
                    "arn:aws:iam::<AWSAccountID>:user/{decube-s3-datalake}"
                ]
            },
            "Action": "kms:Decrypt",
            "Resource": "*"
        }
    ]
}
```

5. `Save Changes`

### Folder partition

* Decube supports ingesting information from multiple dbt projects. You would need to structure the bucket using a format that we define based on the current date.

Given that `base_path` for a single project uses the following format:

* `base_path = ”${year}/${month}/${day}”` where:
  * `year = $(date +%Y)`
  * `month = $(date +%B)`
  * `day = $(date +%d)`
* Example of a folder partition on your S3 - `s3://your-bucket/${base_path}`
  * Where the full path of the folder could be `s3://your-bucket/2024/May/01/`

After setting up the format based on the current date partition, you can proceed to define your own structure.

decube currently supports reading two-level deep bucket structure. You could define how you would want to upload project files into separate directories.

Basically, all of the following are valid bucket path and you can refer to the examples below:

* Assuming the run takes place on the 1st of May 2024:
  * **project\_a**/2024/May/01 - [Example 1](#example-1-multiple-projects)
  * **level1**/2024/May/01 - [Example 1](#example-1-multiple-projects)
  * **dev**/**project\_a**/2024/May/01 - [Example 2](#example-2-multiple-projects-with-environments)
  * **level1**/**level2**/2024/May/01 - [Example 2](#example-2-multiple-projects-with-environments)
  * **the\_project**/2024/May/01 - [Example 3](#example-3-single-project)
  * 2024/May/01 - [Example 4](#example-4-no-project)

### Example 1 - Multiple Projects

* project\_a
  * year=2024
    * month=May
      * day=01
        * \[location of project files]
* project\_b
  * Same as project\_a
* project\_c
  * Same as project\_a

### Example 2 - Multiple Projects with Environments

* dev
  * project\_a
    * year=2024
      * month=May
        * day=01
          * \[location of project files]
  * project\_b
    * Same as project\_a
  * project\_c
    * Same as project\_a
* prod
  * project\_a\_prod
  * project\_b\_prod
  * …

### Example 3 - Single Project

* project\_a
  * year=2024
    * month=May
      * day=01
        * \[location of project files]

### Example 4 - No Project

* year=2024
  * month=May
    * day=01
      * \[location of project files]

### Upload project files

You would need to upload specific files from the `target/` directory into the bucket after your dbt command has concluded.

* `manifest.json`, which is generated by [any command that parses your project](https://docs.getdbt.com/reference/artifacts/manifest-json). Here is an example of a command that generates the file:
  * `dbt run —full-refresh`
    * This [single file](https://docs.getdbt.com/reference/artifacts/manifest-json) contains a full representation of your dbt project's resources (models, tests, macros, etc), including all node configurations and resource properties.
* `run_results.json`, which is generated by a few commands such as `build`, `compile`, and `run` just to name a few (you can refer to the [documentation](https://docs.getdbt.com/reference/artifacts/run-results-json)). Here is an example of a command that generates the file:
  * `dbt build`
    * This [file](https://docs.getdbt.com/reference/artifacts/run-results-json) contains information about a completed invocation of dbt, including timing and status info for each node (model, test, etc) that was executed.
* `catalog.json`, which is only produced by `docs generates` and is **optional**. This is required if you want to acquire column metadata. The command can be run like so:
  * `dbt docs generate`
    * This [file](https://docs.getdbt.com/reference/artifacts/catalog-json) contains information from your [data warehouse](https://docs.getdbt.com/terms/data-warehouse) about the tables and [views](https://docs.getdbt.com/terms/view) produced and defined by the resources in your project.

{% hint style="warning" %}
To ensure the collector runs successfully, you will need to upload in the following manner:

* (in pair) `manifest.json` and `run_results.json` or
* (in triplets) `manifest.json` and `run_results.json` and `catalog.json.`
  {% endhint %}

{% hint style="info" %}
Please be aware in order for the lineage to connect successfully with accuracy, you would need to [configure the source tables](https://docs.getdbt.com/docs/build/sources) on your dbt project.
{% endhint %}

#### Additional Notes

For uploading the project files, you may choose to do the following:

* Only upload the latest project files to the specified bucket where there is only one set of `manifest.json` and `run_results.json` in that bucket for that folder partition at any time.
  * **Caution:** If you were to do it this way, you may lose out information of the runs before the latest project files are processed.
* Retain a series of project files based on the timestamp of when it was run. For example, for each run append a timestamp after the filename:
  * **Do:** manifest\_20240503142827.json
  * **Do not:** 20240503142827\_manifest.json
  * Timestamped project file in this example was generated using the following commands:
    * Using `timestamp=$(date +%Y%m%d%H%M%S)` to create `manifest_${timestamp}.json`

**Note:** To ensure that each project is successfully collected by our metadata collector, we recommend uploading the `manifest.json` and `run_results.json` in the same folder. If you want to include column metadata, make sure you include `catalog.json` as well.

#### Sample Script

Here is a sample script for uploading the project files:

```
#!/bin/bash

# Project name
project_name=some_project

# Generate timestamp
export TZ=UTC
timestamp=$(date +%Y%m%d%H%M%S)

# Generate date-based directory structure
year=$(date +%Y)
month=$(date +%B)
day=$(date +%d)

# Define the base path for S3
base_path="${project_name}/${year}/${month}/${day}"

# Copy project files to S3 with the new structured path
aws s3 cp /path/to/target/manifest.json s3://some-bucket/${base_path}/manifest_${timestamp}.json
aws s3 cp /path/to/target/run_results.json s3://some-bucket/${base_path}/run_results_${timestamp}.json
aws s3 cp /path/to/target/catalog.json s3://some-bucket/${base_path}/catalog_${timestamp}.json
```

You may modify and integrate this into your existing workflows.

## Connecting DBT Core with Decube

After following the above steps, you may start ingesting the metadata from your DBT Core bucket into decube by navigating to `My Account > Data Sources Tab > Connect A New Data Source > DBT Core.`

where 'Path' follows these format:\
s3://some-bucket\
s3://some-bucket/path-to-dbt-core

<figure><img src="/files/ZAQEACr3FOWkCarseGz3" alt=""><figcaption></figcaption></figure>

Please provide the required credentials and click "`Test this connection`" to verify their validity. Afterward, assign a name to your data source, and by selecting the "`Connect This Data Source`" option, your connection between DBT Core and Decube will be successfully established.

### Additional configuration for lineage

Once you have connected your dbt core, you will then need to map the connection sources to the data sources on the decube platform. Refer how to do that in [this documentation](https://docs.decube.io/transformation-tools/additional-configurations).


# Fivetran

Adding Fivetran to your decube connections helps your team to discover, document and monitor the quality of your data transformations.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Data Job, Data Run, Data Task

**Data Quality Monitors**

* Job Failure
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
  {% endtab %}
  {% endtabs %}

Fivetran can map lineage relationships to upstream and downstream objects from the following connectors:

* Upstream Connectors: snowflake, redshift, bigquery, databricks, postgresql, mysql, sql\_server
* Downstream Connectors: snowflake, redshift, bigquery, databricks, postgresql, mysql, sql\_server

## Connection Requirements

### What does the Fivetran connection unlock?

{% hint style="info" %}
This plugin extracts fivetran users, connectors, destinations and sync history. This plugin is in beta and has only been tested on Snowflake connector.
{% endhint %}

### What does the Fivetran connection unlock?

{% embed url="<https://www.loom.com/share/ab76dc0fd3f242abad0a3d38c8555f7d>" %}
Here's what you get when you add a Fivetran connection.
{% endembed %}

### Adding Fivetran connection to decube

From the `My Account` page, select on the Fivetran tile to be brought to the Fivetran connection form.

<figure><img src="/files/tWHBgWzDpUK3a5qd19PJ" alt=""><figcaption><p>Fivetran</p></figcaption></figure>

To connect, you will need to retrieve the API key and API secret from your Fivetran account via these [instructions](https://fivetran.com/docs/rest-api/faq/access-rest-api).

{% hint style="warning" %}
If there is any issues or questions related with connecting your data sources, please reach out to us via our Live chat for support.
{% endhint %}


# Airflow

Adding Airflow to your decube connections helps your team to discover, document and monitor the quality of your pipelines.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Data Job, Data Run, Data Task

**Data Quality Monitors**

* Job Failure
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
  {% endtab %}
  {% endtabs %}

## Connection Requirements

* Airflow Username
* Airflow User Password
* Airflow API host address
* Airflow API enabled and set to Basic Auth. See [Airflow Documentation](https://airflow.apache.org/docs/apache-airflow/stable/administration-and-deployment/security/api.html#basic-authentication) for this.
* Airflow Version 2.3.0 and above (Version >= 2.0.0 and < 2.3.0 may not work fully)

<figure><img src="/files/yqzfv7rI92edPfYyJSH5" alt=""><figcaption><p>Airflow</p></figcaption></figure>

## Creating an Airflow User for Decube

1. Go to Security > List Users

<figure><img src="/files/tosNnLk3fsFv8Uc9Nphp" alt=""><figcaption></figcaption></figure>

2. Click on `"+"` to Add User

<figure><img src="/files/Yro1CT33nSa7HtS23meJ" alt=""><figcaption></figcaption></figure>

3. Insert information for new decube user
   1. Username - suggested value: `decube`
   2. Email - <collectors@decube.io>
   3. Role - Minimum required `Op` (which is a default role from Airflow)
   4. Password - Use a strong password

<figure><img src="/files/tRo1FDRwUXGdvPpaUmu4" alt=""><figcaption></figcaption></figure>

## Airflow API is not Publicly Accessible

* For decube to monitor your Airflow service, we will require that the Airflow API be publicly accessible or privately accessible to a SSH bastion host. Instruction on setting up a bastion host can be found here [SSH Tunneling](/security-and-connectivity/ssh-tunneling)


# AWS Glue

View catalogued assets within your AWS Glue, or leverage AWS Athena to add data observability capabilities and monitor Iceberg tables.

{% hint style="info" %}
**Glue + Athena architecture:** AWS Glue handles metadata extraction and lineage (including Iceberg tables). Amazon Athena is the optional compute layer that enables profiling, data preview, data quality monitors, and Iceberg table observability. Without Athena, this connector provides structural metadata only. See [Enable Athena for Data Observability](#enable-athena-for-data-observability) to set it up.
{% endhint %}

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**Without Athena**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column, Data Job, Data Run, Data Task
* **Configurable Collection** — selective ingestion of schemas/workspaces in Data Source Management

**With Athena enabled**

* **Profiling** — data profiling on the Profiler tab
* **Preview** — sample data preview
* **Data Quality** — data quality monitoring and observability
* **Iceberg table support** — profiling, preview, and data quality monitors on Iceberg tables

**Data Quality Monitors (with Athena enabled)**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Job Failure
  {% endtab %}
  {% endtabs %}

## Minimum Requirement

To connect your AWS Glue to decube, we will need the following information:

Choose authentication method:

a. [**AWS Identity**](#a.-aws-role):

* Select AWS Identity
* Customer AWS Role ARN
* Region
* Enable Athena (Optional) - Read more in this [section](#enable-athena-for-data-observability). If Athena is enabled,
  * Workgroup
  * Bucket Name
* Data source name

b. **AWS Access Key**:

* Access Key ID
* Secret Access Key
* Region
* Enable Athena (Optional) - Read more in this [section](#enable-athena-for-data-observability). If Athena is enabled,
  * Workgroup
  * Bucket Name
* Data source name

<figure><img src="/files/ibYX9QV9lOG32R6y2t5X" alt=""><figcaption><p>AWS Glue</p></figcaption></figure>

## Connection Options:

### a. AWS Role

{% hint style="info" %}
This section will create a Customer AWS Role within your AWS account that has the right set of permission to access your data sources.
{% endhint %}

* Step 1: Go to your AWS Account → IAM Module → Roles
* Step 2: Click on **Create role**

<figure><img src="/files/oqh7ru346tgegU5ag3yO" alt=""><figcaption></figcaption></figure>

* Step 3: Choose **Custom trust policy**

<figure><img src="/files/69oNLlXEpHGoPBgwyZUV" alt=""><figcaption></figcaption></figure>

* Step 4: Specify the following as the trust policy, replacing `DECUBE-AWS-IDENTITY-ARN` and `EXTERNAL-ID` with values from [AWS Identities](/security-and-connectivity/aws-identities#generating-a-decube-aws-identity)

```
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "AWS": "<DECUBE-AWS-IDENTITY-ARN>"
            },
            "Action": "sts:AssumeRole",
            "Condition": {
                "StringEquals": {
                    "sts:ExternalId": "<EXTERNAL-ID>"
                }
            }
        }
    ]
}
```

* Step 5: Click next to proceed to attach policy
* Step 6: Click on **Attach policies directly** and search for `AWSGlueServiceRole` and add this policy

<figure><img src="/files/fXXDEDJOev02oHZ1FXHg" alt=""><figcaption></figcaption></figure>

* Step 6: Click next and specify a role name. For this documentation, the name will be presumed to be CustomerAWSRole but can be set to any value.

### b. AWS IAM User

* Step 1: Login to AWS Console and proceed to IAM > User > Create User

<figure><img src="/files/OQfuTRN6wnAeEAsVR6kZ" alt=""><figcaption></figcaption></figure>

* Step 2: Click on attach policies directly and search for `AWSGlueServiceRole`

<figure><img src="/files/YB2nJTHsQ2CfocAUrwRx" alt=""><figcaption></figcaption></figure>

* Step 3: Review and create your user

<figure><img src="/files/fd1NZxJ5D4l1k9uKnyaC" alt=""><figcaption></figcaption></figure>

* Step 4: Navigate to the newly created user and click on `Create access key`

<figure><img src="/files/esOafPde94OPokkJe01f" alt=""><figcaption></figcaption></figure>

* Step 5: Choose `Application running outside AWS`

<figure><img src="/files/I11zud2urGiLpjjEhiHn" alt=""><figcaption></figcaption></figure>

* Step 6: Save the provided access key and secret access key. You will not be able to retrieve these keys again

<figure><img src="/files/MeoD0UaxqzIePBJzSTtN" alt=""><figcaption></figcaption></figure>

## Enable Athena for Data Observability

{% hint style="info" %}
This section is applicable if you intend to enable monitors on your AWS Glue source. This includes monitoring on Iceberg tables as AWS Athena will be required to query Iceberg tables.
{% endhint %}

AWS Glue, by itself, does not provide native support for Data Quality Monitoring. To address this, we leverage AWS Athena, a serverless, interactive query service, to analyze and query the data that was produced by glue and stored in AWS S3. Because of that, Decube requires additional policies to be attached to the IAM user created in this step [AWS Glue | Decube](#aws-iam-user).

### Configuring AWS Athena

You will need to set up these items:

1. Creating an s3 bucket to store Athena query results.
2. Creating an Athena Workgroup
3. Optional - Athena Data Source name

Athena saves the results of queries in an s3 bucket. The location of the bucket will then be attached in one of the policies of the next section. Athena Workgroup is required for some of the policies that we will attach to the IAM user as well.

### Creating an s3 bucket

1. Go to `S3` > `Bucket` and click on `Create bucket`
2. For bucket name, we suggest `decube-athena-query-results`.
3. For `Object Ownership`, select `ACLs disabled`.
4. Click on `Create bucket`.
5. Take note of the ARN for the bucket, we will refer it as `decube-athena-query-results` in the following sections when setting up Athena.

### Creating an Athena Workgroup

1. Go to `Amazon Athena` > `Administration` > `Workgroups`
2. Click on `Create Workgroup`

<figure><img src="/files/nOyw59HqZgg7dYZLeu0c" alt=""><figcaption><p>Create workgroup details.</p></figcaption></figure>

3. Fill in Workgroup name. Recommended name here is: `decube-athena-workgroup`.
4. Select `Athena SQL` as `Analytics engine`.
5. Select `Manual` for `Upgrade query engine`.
6. Select `Athena engine version 3` as `Query engine version`.

<figure><img src="/files/71QN8X9NlVMdxEnnaOL4" alt=""><figcaption><p>Select AWS IAM and add location of query result.</p></figcaption></figure>

7. For Authentication, select `AWS Identity and Access Management (IAM)`.
8. For `Query result configuration`, specifically `Location of query result` **,** fill in the location of the bucket from the previous section, if you’re following the name convention it would be `s3://decube-athena-query-results` .
9. Click on `Create workgroup`.
10. Take note of the name of the workgroup.

### Adding policies to IAM User

{% hint style="warning" %}
Ensure that step in [previous section](#aws-iam-user) to set up IAM User has been completed first before this section.
{% endhint %}

1. Go to IAM > Users and search for the user previously created for Decube to ingest Glue and click on it.
2. On the `Permissions` tab, click on `Add permissions` > `Create inline policy`.

<figure><img src="/files/3Mk4kBgFztGbAN7mE0id" alt=""><figcaption><p>Go to Create inline policy.</p></figcaption></figure>

3. On the `Policy editor` tab, click on `JSON.`Click `Next`.

<figure><img src="/files/6hLbozOYdRj2NA8s44m1" alt=""><figcaption><p>Select JSON</p></figcaption></figure>

4. Copy and paste these policies onto the form provided, **note to change the block on Resource** accordingly.

```json
{
	"Version": "2012-10-17",
	"Statement": [
		{
			"Sid": "DecubeAthenaS3Ingest",
			"Effect": "Allow",
			"Action": [
				"s3:GetObject",
				"s3:GetBucketLocation",
				"athena:GetTableMetadata",
				"athena:StartQueryExecution",
				"athena:GetQueryResults",
				"athena:GetDatabase",
				"athena:GetDataCatalog",
				"athena:ListQueryExecutions",
				"athena:GetWorkGroup",
				"athena:StopQueryExecution",
				"athena:GetQueryResultsStream",
				"athena:ListDatabases",
				"athena:GetQueryExecution",
				"athena:ListTableMetadata",
				"athena:BatchGetQueryExecution"
			],
			"Resource": [
				"arn:aws:athena:*:{account id}:datacatalog/{specify a data catalog or *}",
				"arn:aws:athena:*:{account id}:workgroup/{workgroup_name or *}",
			  // example
				// "arn:aws:athena:*:1234567:datacatalog/*",
				// "arn:aws:athena:*:1234567:workgroup/decube-athena-workgroup ",
				
				// example - all buckets to be monitored by Athena
				// "arn:aws:s3:::decube-glue_results/*",
				// "arn:aws:s3:::decube-glue_results"
			]
		},
		{
			"Sid": "DecubeS3AthenaOutput",
			"Effect": "Allow",
			"Action": [
				"s3:PutObject",
				"s3:GetObject",
				"s3:ListBucketMultipartUploads",
				"s3:AbortMultipartUpload",
				"s3:ListBucket",
				"s3:GetBucketLocation",
				"s3:ListMultipartUploadParts"
			],
			"Resource": [
				// example. ARN from athena input bucket
				// "arn:aws:s3:::decube-athena-query-results/*",
				// "arn:aws:s3:::decube-athena-query-results"
			]
		}
	]
}
```

5. Click Next . We recommend naming the policy decube-athena-s3. Finally click on Create policy.

<figure><img src="/files/vzsKsWp8pIKVCWaGjoI7" alt=""><figcaption><p>Create policy name.</p></figcaption></figure>

## OpenLineage with AWS Glue

This section is applicable if you intend to view lineages from your AWS Glue jobs. OpenLineage is an open framework for data lineage collection and analysis. At its core is an extensible specification that systems can use to interoperate with lineage metadata.

Follow below steps to [enable OpenLineage on AWS Glue](https://openlineage.io/docs/integrations/spark/quickstart/quickstart_glue/):

1. **Specify the OpenLineage JAR URL**

* In the **Job details** tab, navigate to **Advanced properties** → **Libraries** → **Dependent Jars path**

<figure><img src="/files/EeUSlU4BbLYp3gZ7TLLd" alt=""><figcaption></figcaption></figure>

* Use the URL directly from [**Maven Central openlineage-spark**](https://mvnrepository.com/artifact/io.openlineage/openlineage-spark)
* **`Ensure you select the version for Scala 2.12, as Glue Spark is compiled with Scala 2.12, and version 2.13 won't be compatible.`**
* On the page, for the specific OpenLineage version for Scala 2.12, copy the URL of the jar file from the Files row and use it in Glue.
* **Alternatively**, upload the jar to an **S3 bucket** and use its URL. The URL should use the `s3` scheme: `s3://<your bucket>/path/to/openlineage-spark_2.12-<version>.jar`

2. **Add OpenLineage configuration in Job Parameters**

   In the same **Job details** tab, add a new property under **Job parameters**:

   * Use the format **`param1=value1 --conf param2=value2 ... --conf paramN=valueN`**.
   * Make sure every parameter except the first has an extra **`--conf`** in front of it.
   * Example: `spark.extraListeners=io.openlineage.spark.agent.OpenLineageSparkListener --conf spark.openlineage.transport.type=http --conf spark.openlineage.transport.url=https://integrations.<Region>.decube.io --conf spark.openlineage.transport.endpoint=/integrations/openlineage/webhook/<webhook-uuid> --conf spark.openlineage.transport.auth.type=api_key --conf spark.openlineage.transport.auth.apiKey=<webhook-key>`
3. **Set User Jars First Parameter**

* Add the --user-jars-first parameter and set its value to true

<figure><img src="/files/6IX2whTdGsNn0OD8EZf6" alt=""><figcaption></figcaption></figure>

## Verification

* To confirm that OpenLineage registration has been successful, check the logs for the following entry:

```
INFO SparkContext: Registered listener io.openlineage.spark.agent.OpenLineage
SparkListener
```

* If you see this log message, it indicates that OpenLineage has been correctly registered with your AWS Glue job.

7. Insert the "access key" and "secret key" with "region" of the connection form, then test the connection. If it is successful, you can now add the name and connect to the data source.


# Azure Data Factory

Azure Data Factory is a cloud-based data integration and orchestration service by Microsoft Azure. See all your ETLs in decube and monitor the status of your jobs.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Data Job, Data Run, Data Task
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure
  {% endtab %}
  {% endtabs %}

Azure Data Factory can map lineage relationships to upstream and downstream objects from the following connectors:

* Upstream Connectors: postgresql, mysql, synapse, azure\_server, databricks, sql\_server, redshift, bigquery, adls
* Downstream Connectors: postgresql, mysql, synapse, azure\_server, databricks, sql\_server, redshift, bigquery, adls

## Connection Requirements

From our Azure account, we will need the following information:

* Tenant ID
* Client ID
* Client Secret
* Subscription ID
* Resource Group Name
* Factory Name
* Data Source Name

<figure><img src="/files/rTB8vNdkpJV3Puq1xDLL" alt=""><figcaption><p>Azure Data Factory</p></figcaption></figure>

### How to connect

1. On the Azure Home Page, go to `Azure Active Directory`.

<figure><img src="/files/fYWKXkVNOEMTTsgmPZa6" alt=""><figcaption></figcaption></figure>

2. Go to `App registrations`.

<figure><img src="/files/AYbMSbCP9LwS3JqMsMz2" alt=""><figcaption></figcaption></figure>

3. Click on `New registration`.

<figure><img src="/files/CvSZIIAC61TodJPFCFJB" alt=""><figcaption></figcaption></figure>

4. Click `Register`.

<figure><img src="/files/iiqIxFcrbHBFGZtzS3rb" alt=""><figcaption></figcaption></figure>

5. Save the `Application (client) ID` and `Directory (tenant) ID`.
6. Click `Add a certificate or secret`.
7. Go to `Client secrets` and click `+ New client secret`.

<figure><img src="/files/SzwLmJjq5BI3kkYzsNW8" alt=""><figcaption></figcaption></figure>

8. Click `Add`.

<figure><img src="/files/LkeMT798jDzvVUprbQd8" alt=""><figcaption></figcaption></figure>

9. Copy and save the `Value`.

<figure><img src="/files/YwbsKAOB1VptaFW7yWSL" alt=""><figcaption></figcaption></figure>

10. Go to Data Factories and click the factory you wanna add.

<figure><img src="/files/p1nI4M3vbPzNPbijszQY" alt=""><figcaption></figcaption></figure>

11. Go to `Access Control (IAM)`.

<figure><img src="/files/ZTQIJU03NKVfg7xXQz29" alt=""><figcaption></figcaption></figure>

12. Click on `+ Add` and select `Add role assignment`.

<figure><img src="/files/IMTTxF3q6ScWyGb7ADY0" alt=""><figcaption></figcaption></figure>

13. Select `Data Factory Contributor`.

<figure><img src="/files/Mt5ngQ1Bso8tAOVhhfmV" alt=""><figcaption></figcaption></figure>

14. Go to `Members` tab and Click on `+ Select members`.

<figure><img src="/files/uPbVykoymIFFEvHArM8S" alt=""><figcaption></figcaption></figure>

15. Search and select the service principal that was created in the previous step. Click on `Select`.

<figure><img src="/files/Wp3rtUACk9RvDDErp4Ck" alt=""><figcaption></figcaption></figure>

16. Go to `Review + assign` tab and Click `Review + assign`

<figure><img src="/files/ucDLmh4UgqL0WwD58TQZ" alt=""><figcaption></figcaption></figure>

17. Go to Data Factories and select the factory you wanna add, copy the `Name` and `Resource group`.

<figure><img src="/files/pELdq8DFoHJjkeNcJwB8" alt=""><figcaption></figcaption></figure>

18. Copy the `Subscription ID`.

<figure><img src="/files/RiqnkhsUIe05H09WdIWr" alt=""><figcaption></figcaption></figure>

19. Fill all the required fields in the connection form, and click on Test this connection once connection is successful, give your database a name and connect the data source.


# Apache Spark

See lineages from spark jobs in Decube Catalog.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Data Job, Data Run, Data Task

**Data Quality Monitors**

* Job Failure
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
  {% endtab %}
  {% endtabs %}

Apache Spark can map lineage relationships to upstream and downstream objects from the following connectors:

* Upstream Connectors: postgresql, adls
* Downstream Connectors: postgresql, adls

## Connection Requirements

Please see the instructions and minimum requirements for configuration in each data source below:

* [Azure Synapse](/transformation-tools/apache-spark/apache-spark-in-azure-synapse)


# Apache Spark in Azure Synapse

Connecting Decube to Apache Spark specifically for Azure Synapse.

This guide outlines the steps necessary to integrate Decube with an Apache Spark instance in Azure Synapse Analytics Workspace. The process is divided into two primary sections: configuring the Decube platform and setting up the necessary components within the Azure Synapse Analytics Workspace.

## **Minimum Requirements: Credentials and Access**

* Before beginning the setup, ensure that you have the following credentials and access:
* **Decube Data Source:**
  * Name of the Source
* **Client - Azure Synapse Workspace:**
  * Must have **WRITE** access to Azure Synapse Workspace Libraries.
* **Required roles include:**
  * Synapse Administrator
  * Synapse Apache Spark Administrator
  * Synapse Contributor
  * Synapse Artifact Publisher

## **How to Connect**

The connection process is divided into two main parts:

* *Decube* platform configuration
* *Azure Synapse Analytics Workspace* setup

**Part 1: Decube Platform Configuration**

**1. Create Spark Data Source**

* Input the name of the connector. This will be the identifier for your data source within Decube.
* You can include exclusion filters which exclude specific tables and lineage paths from being ingested.

<figure><img src="/files/3Ni2NIv67Dh25DEHK0M9" alt=""><figcaption></figcaption></figure>

**2. Copy Credentials**

* After creating the data source, Decube will provide you with credentials. These credentials are essential for the integration process and will be used later in the Azure Synapse setup.

<figure><img src="/files/p4gnAWvcP2EvJx6YRwxb" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
**Important:** Keep these credentials secure as they are required for the integration with the Azure Synapse Apache Spark instance.
{% endhint %}

**Part 2: Azure Synapse Analytics Workspace Setup**

1. **Setup OpenLineage as a Workspace Package**

To track and manage data lineage within Spark jobs, the OpenLineage library must be added as a workspace-wide package, so we can add to the *Apache Spark Pool* as packages later.

**Steps:**

**Download the OpenLineage Binary:**

Navigate to the <https://mvnrepository.com/artifact/io.openlineage/openlineage-spark> and download the appropriate version of the OpenLineage binary. For most use cases, the version 1.20.5 - Scala 2.13 is recommended.

{% hint style="info" %}
Make sure that the Scala Versions match for Apache Spark and OpenLineage Library.

The versions before `1.8.0` does not have support of `Scala 2.13` , only `Scala 2.12` is recommended.
{% endhint %}

* **Below is the example how to download the OpenLineage Binary from Maven:**
* After choosing the version `1.20.5 - Scala 2.13`

<figure><img src="/files/vpE9M89aRHHvXtaEEAAm" alt=""><figcaption></figcaption></figure>

* Click on the `jar` label to download the binary.

<figure><img src="/files/QCb57Z9dmz5TzaSCO5Ga" alt=""><figcaption></figcaption></figure>

**Upload the Binary to Azure Synapse:**

* Go to Synapse Studio.
* In the left panel, navigate to **Manage** > **Configurations + Libraries** > **Workspace Packages** > **Upload**.
* Upload the OpenLineage jar binary from **Download the OpenLineage Binary.**

<figure><img src="/files/4nresldA20JJ2cI18zTw" alt=""><figcaption></figcaption></figure>

2. **Install OpenLineage in Apache Spark Pool Environment**

Now that OpenLineage is a workspace package, it needs to be installed in the default Apache Spark Pool environment.

**Steps:**

1. In Synapse Studio, go to **Manage** > **Apache Spark Pools**.
2. Select the pool to configure and navigate to **Packages**.
3. Choose the OpenLineage jar uploaded in the previous step and install it.

<figure><img src="/files/XUvvJPBeHSCfhnkrvIJ5" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/12rabYCkI6mqIzN9tuW0" alt=""><figcaption></figcaption></figure>

* Select the `OpenLineage` jar that was uploaded as workspace packages.

{% hint style="info" %}
**Note:** The installation process may take some time as the package is being integrated into the Spark environment.
{% endhint %}

**3. Configure Apache Spark with Decube Credentials**

To finalize the setup, you need to configure Apache Spark with the necessary settings to communicate with Decube.

* We will setup the Apache Spark Configurations for `OpenLineage` , by using the credentials copied during **Copy Credentials** portion of Data Source setup, it is to be noted that the credentials can also be seen after setting up in the **Modify** screen.

**Configuration Parameters:**

```jsx
spark.extraListeners io.openlineage.spark.agent.OpenLineageSparkListener
spark.openlineage.transport.type http
spark.openlineage.transport.url https://integrations.<Region>.decube.io
spark.openlineage.transport.auth.type api_key
spark.openlineage.transport.auth.apiKey <api-key>
spark.openlineage.transport.endpoint /integrations/openlineage/webhook/<webhook-uuid>
spark.openlineage.namespace [Namespace]
```

**Api-key**: The API key copied from the Decube platform.

**Webhook-uuid**: The webhook UUID copied from the Decube platform.

**\[Namespace]**: A custom namespace determined by the client.

{% hint style="info" %}
For full reference on how to add the configurations above, see <https://learn.microsoft.com/en-us/azure/synapse-analytics/spark/apache-spark-azure-create-spark-configuration>
{% endhint %}

4. **Add to Existing Spark Pool Configuration**

<figure><img src="/files/uQo6TRr7gku9QaQdcwWK" alt=""><figcaption></figcaption></figure>

* **Input the configurations provided in** [**Apache Spark**](https://docs.decube.io/transformation-tools/apache-spark) **with the placeholders filled with the values required.**

For Example:

<figure><img src="/files/6sGYr6Y0Ul9Pg9xSBbVn" alt=""><figcaption></figcaption></figure>

### Expected output

Once your Spark has been successfully set up, you should be able to see the Data Jobs in the Catalog (which are named after app name in Spark Config).

<figure><img src="/files/8WlWod3VFT4KbAvSQWvn" alt=""><figcaption><p>Example Data Jobs in Catalog</p></figcaption></figure>

You will also be able to see lineages extracted from the workflow.

<figure><img src="/files/IpMNkp7bk4VM4Oubdbuv" alt=""><figcaption><p>Example of 2 csv tables joined onto a parquet file in ADLS.</p></figcaption></figure>

## Exclusion Filters

Exclusion filters let you exclude specific tables and lineage paths from being ingested by Decube. This is useful when your Spark jobs produce metadata for staging tables, temporary paths, or other assets you do not want tracked in your catalog.

You configure exclusion filters directly in the Decube UI on your Apache Spark data source settings page. Each filter expects a Python RegEx-compliant regular expression for its fields.

### Supported Filter Types

{% tabs %}
{% tab title="ADLS Gen2 Path" %}
Matches tables with an ADLS Gen2 URI in the format:

```
abfss://<container-name>@<service-name>.dfs.core.windows.net/<path>
```

| Field          | Description                        |
| -------------- | ---------------------------------- |
| container-name | The ADLS container name            |
| service-name   | The ADLS storage account name      |
| path           | The file path within the container |

**Example** — exclude everything under the `discard/` path in the `decube` container across all storage accounts:

| Field          | Value        |
| -------------- | ------------ |
| container-name | `decube`     |
| service-name   | `.*`         |
| path           | `discard/.*` |

This excludes `abfss://decube@test.dfs.core.windows.net/discard/test/file` but not `abfss://decube@test.dfs.core.windows.net/nodiscard/test/file`.
{% endtab %}

{% tab title="Snowflake" %}
Matches Snowflake tables by their fully qualified identifier.

| Field              | Description                                                           |
| ------------------ | --------------------------------------------------------------------- |
| account-identifier | Your Snowflake account in `<organization-name>-<account-name>` format |
| database           | The database name                                                     |
| schema             | The schema name                                                       |
| table              | The table name                                                        |

**Example** — exclude all tables in the `TEST` schema of `WORKDATABASE` on the `decube-test` account:

| Field              | Value          |
| ------------------ | -------------- |
| account-identifier | `decube-test`  |
| database           | `WORKDATABASE` |
| schema             | `TEST`         |
| table              | `.*`           |

This excludes `snowflake://decube-test/WORKDATABASE.TEST.TABLE` but not `snowflake://decube-test/PROD.TEST.TABLE`.
{% endtab %}

{% tab title="S3" %}
Matches tables stored in Amazon S3.

| Field       | Description                             |
| ----------- | --------------------------------------- |
| bucket-name | The S3 bucket name                      |
| object-key  | The object key (path) within the bucket |

**Example** — exclude all CSV files under `raw/data/` in the `decube-datalake` bucket:

| Field       | Value              |
| ----------- | ------------------ |
| bucket-name | `decube-datalake`  |
| object-key  | `raw/data/.*\.csv` |

This excludes `s3://decube-datalake/raw/data/report.csv` but not `s3://decube-datalake/raw/data/report.json` or `s3://decube-datalake/bronze/report.csv`.
{% endtab %}

{% tab title="Generic Regex" %}
Use the generic regex filter when the table format does not match any of the specific filter types above. The regex is matched against the full table identifier regardless of type.

| Field | Description                                              |
| ----- | -------------------------------------------------------- |
| regex | A regular expression matched against the full table path |

**Example** — exclude any table whose path contains `/ignored/`:

| Field         | Value           |
| ------------- | --------------- |
| regex         | `.*/ignored/.*` |
| {% endtab %}  |                 |
| {% endtabs %} |                 |

{% hint style="info" %}
You can add multiple exclusion filters of different types on the same data source. Each filter is evaluated independently — a table is excluded if it matches any filter.
{% endhint %}


# OpenLineage (BETA)

This document provides a step-by-step guide to connecting with the OpenLineage connector and viewing lineage data from jobs using the OpenLineage framework.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Data Job, Data Task, Data Run

**Data Quality Monitors**

* Job Failure
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure

**Data Quality Monitors**

* Freshness
* Volume
* Field Health
* Custom SQL
* Schema Drift
  {% endtab %}
  {% endtabs %}

## Connection Requirements

* Step 1: Go to **My Account** and click on the **Integrations** tab
* Step 2: Go to the **Connect a new data source** section
* Step 3: Click on the **OpenLineage** icon.
* Step 4: Enter a **name** for the data source and click **Submit.**

The exclusion filters can be added to exclude specific tables and lineage paths from being ingested. [See more here](#exclusion-filters).

<figure><img src="/files/XtGoUs3h1mmOHsvyi1Il" alt=""><figcaption></figcaption></figure>

Step 5: A **Webhook UUID** and an **API Key** will be provided. **Copy** them into your connector’s configuration settings.

<figure><img src="/files/aiW16XpomD7t8U1Rj5m7" alt=""><figcaption></figcaption></figure>

## Webhook Endpoint

Payload must submitted to the following endpoint:

```
https://integrations.<region>.decube.io/integrations/openlineage/webhook/<webhook-uuid>
```

## Submitting Payload to OpenLineage Webhook

If you're using these tools, please follow the respective documentation in the OpenLineage website.

| Tool         | Documentation                                                                                                                    |
| ------------ | -------------------------------------------------------------------------------------------------------------------------------- |
| Airflow      | [openlineage.io/docs/integrations/airflow/usage](https://openlineage.io/docs/integrations/airflow/usage)                         |
| Apache Spark | [openlineage.io/docs/integrations/spark/configuration/usage](https://openlineage.io/docs/integrations/spark/configuration/usage) |
| Apache Flink | [openlineage.io/docs/integrations/flink/configuration](https://openlineage.io/docs/integrations/flink/configuration)             |

## Custom Integration

If you want to create your own integration for your tools, follow these steps:

* Submit the webhook payload to the above endpoint.
* Use the Bearer token system for authentication.

Example request:

```
curl -X POST \
   -H "Authorization: Bearer <api-key>" \
   -H "Content-Type: application/json" \
   --data '{}' \
   https://integrations.<region>.decube.io/integrations/openlineage/webhook/<webhook-uuid>
```

## Exclusion Filters

Exclusion filters let you exclude specific tables and lineage paths from being ingested by Decube. This is useful when your OpenLineage jobs produce metadata for staging tables, temporary paths, or other assets you do not want tracked in your catalog.

You configure exclusion filters directly in the Decube UI on your OpenLineage data source settings page. Each filter expects a Python RegEx-compliant regular expression for its fields.

### Supported Filter Types

{% tabs %}
{% tab title="ADLS Gen2 Path" %}
Matches tables with an ADLS Gen2 URI in the format:

```
abfss://<container-name>@<service-name>.dfs.core.windows.net/<path>
```

| Field          | Description                        |
| -------------- | ---------------------------------- |
| container-name | The ADLS container name            |
| service-name   | The ADLS storage account name      |
| path           | The file path within the container |

**Example** — exclude everything under the `discard/` path in the `decube` container across all storage accounts:

| Field          | Value        |
| -------------- | ------------ |
| container-name | `decube`     |
| service-name   | `.*`         |
| path           | `discard/.*` |

This excludes `abfss://decube@test.dfs.core.windows.net/discard/test/file` but not `abfss://decube@test.dfs.core.windows.net/nodiscard/test/file`.
{% endtab %}

{% tab title="Snowflake" %}
Matches Snowflake tables by their fully qualified identifier.

| Field              | Description                                                           |
| ------------------ | --------------------------------------------------------------------- |
| account-identifier | Your Snowflake account in `<organization-name>-<account-name>` format |
| database           | The database name                                                     |
| schema             | The schema name                                                       |
| table              | The table name                                                        |

**Example** — exclude all tables in the `TEST` schema of `WORKDATABASE` on the `decube-test` account:

| Field              | Value          |
| ------------------ | -------------- |
| account-identifier | `decube-test`  |
| database           | `WORKDATABASE` |
| schema             | `TEST`         |
| table              | `.*`           |

This excludes `snowflake://decube-test/WORKDATABASE.TEST.TABLE` but not `snowflake://decube-test/PROD.TEST.TABLE`.
{% endtab %}

{% tab title="S3" %}
Matches tables stored in Amazon S3.

| Field       | Description                             |
| ----------- | --------------------------------------- |
| bucket-name | The S3 bucket name                      |
| object-key  | The object key (path) within the bucket |

**Example** — exclude all CSV files under `raw/data/` in the `decube-datalake` bucket:

| Field       | Value              |
| ----------- | ------------------ |
| bucket-name | `decube-datalake`  |
| object-key  | `raw/data/.*\.csv` |

This excludes `s3://decube-datalake/raw/data/report.csv` but not `s3://decube-datalake/raw/data/report.json` or `s3://decube-datalake/bronze/report.csv`.
{% endtab %}

{% tab title="Generic Regex" %}
Use the generic regex filter when the table format does not match any of the specific filter types above. The regex is matched against the full table identifier regardless of type.

| Field | Description                                              |
| ----- | -------------------------------------------------------- |
| regex | A regular expression matched against the full table path |

**Example** — exclude any table whose path contains `/ignored/`:

| Field         | Value           |
| ------------- | --------------- |
| regex         | `.*/ignored/.*` |
| {% endtab %}  |                 |
| {% endtabs %} |                 |

{% hint style="info" %}
You can add multiple exclusion filters of different types on the same data source. Each filter is evaluated independently — a table is excluded if it matches any filter.
{% endhint %}


# Additional configurations

Here's what you can do with to complete the lineage of your connections.

To identify the lineage across your tools from origin to destination, you will need to connect your data sources to decube. For example, if you want to see a transformation from PostgreSQL table to dbt, you'll need to have both PostgreSQL and dbt connected to see the full lineage.

{% hint style="info" %}
It's best if you connect your data warehouses or relational databases first to decube if you haven't done so, before connecting your transformation tools or BI tools.
{% endhint %}

From the tools you've connected, we figure out the connection names referenced within your queries and list them out in the **Additional Config** so you can map them to existing connections within decube.

### Mapping the data source connections

1. Once you've connected your connections, all you need to do now is to head to the **My Account** page and see all your connections in the Data Sources tab.

<figure><img src="/files/pjtjRTDoAqWOJVqPu7iT" alt=""><figcaption><p>Head to the My Account page to see all the connected Data Sources</p></figcaption></figure>

2. Click on the **Modify** button on your data source. You will land on the **Credentials** tab.

<figure><img src="/files/PhMagr3sZnUdNyqp3Z2w" alt=""><figcaption><p>Modifying Data Source</p></figcaption></figure>

3. Go ahead and map the connections sources to the data sources that you have added into decube. Once done, click on **Save**.


# Amazon Athena

Use Amazon Athena as Decube's compute engine for AWS Glue data sources, enabling profiling, data preview, data quality monitors, and Iceberg table observability.

Amazon Athena is supported in Decube as a **compute engine** paired with the [AWS Glue](/transformation-tools/aws-glue). It is not a standalone catalog source — Athena works alongside Glue to execute queries against your S3-backed tables, unlocking data observability capabilities that Glue alone cannot provide.

## How the Glue + Athena pairing works

AWS Glue serves as the metadata catalog: it holds your table schemas, column definitions, and data job lineage (including Iceberg tables). Decube connects to Glue to extract and display this structural metadata.

Athena is a serverless query engine that runs SQL against data stored in Amazon S3. Because Glue does not natively support query execution, Decube uses Athena to power profiling runs, sample data previews, and data quality monitors. When you enable Athena on a Glue connection, Decube routes compute operations through Athena while continuing to use Glue for all metadata and lineage.

## What becomes available with Athena enabled

Without Athena, a Glue connection provides metadata and lineage only. Enabling Athena unlocks:

* **Profiling** — run data profiles on your Glue-catalogued tables from the Profiler tab
* **Data preview** — view sample rows directly in the Catalog
* **Data quality monitors** — set up and run all five monitor types:
  * Freshness
  * Volume
  * Field Health
  * Custom SQL
  * Schema Drift
* **Iceberg table support** — profiling, data preview, and all data quality monitors on Iceberg tables (Athena is required to query Iceberg format)

{% hint style="info" %}
Job Failure monitors are not supported on Glue sources regardless of whether Athena is enabled.
{% endhint %}

## Setting up Athena

Athena is configured as part of the AWS Glue connector setup. You will need:

* An S3 bucket to store Athena query results
* An Athena Workgroup
* Additional IAM policies attached to the Glue IAM user

Full setup instructions are in the [AWS Glue](/transformation-tools/aws-glue#enable-athena-for-data-observability) section of the AWS Glue connector page.


# Salesforce

Adding Salesforce to your decube connections helps your team find relevant datasets and apply governance policies via the data catalog

{% hint style="info" %}
This connector will be released to BETA soon. Stay tuned.
{% endhint %}


# Tableau

Add a Tableau connection to your decube so that you can discover the lineage of your assets from source to dashboards.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Chart, Dashboard
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure
  {% endtab %}
  {% endtabs %}

Tableau supports lineage mapping from the following sources:

* Upstream Connectors: bigquery, snowflake, mysql, redshift, postgresql

## Connection Requirements

### Prerequisite - Enable Metadata API

To enable successful metadata ingestion for Tableau so you can see Dashboards and Charts from Tableau, you will need to enable the Tableau Metadata API for Tableau Server. For the guide from Tableau, refer to [this](https://help.tableau.com/current/api/metadata_api/en-us/docs/meta_api_start.html#enable-the-tableau-metadata-api-for-tableau-server).

### Connecting to decube

We require four important information to connect a Tableau instance. For two of those, head to your Tableau dashboard and take a look at your browser's search bar. Grab the URL and it would look like something below:

```
https://<example-server-url.com>/#/site/<example-site-id>/projects
```

Here we need `https://<example-server-url.com>` as `Server URL` and `<example-site-id>` as the `Site ID`

Get a `Personal Access Token` following [this step](#getting-a-personal-access-token).

Then we can fill the connections as so, remember to replace the example URL with your own.

* `Server URL`: \<example-server-url.com>
* `Site ID`: \<example-site-id>

{% hint style="info" %}
If you are using Tableau Server (On-premise), you can leave the Site ID field empty upon submission.
{% endhint %}

* `Token Name`: Personal Access Token Name
* `Token Secret`: Personal Access Token Secret

<figure><img src="/files/VHGoeFJJWi9jPIgbkHk2" alt=""><figcaption><p>Tableau</p></figcaption></figure>

The `Source Name` will be for you to differentiate and recognize particular sources within the decube application.

### Getting a Personal Access Token

decube prefers connecting to Tableau via a Personal Access Token. It is recommended to create a new Access Token for decube and ensure that it is kept safe and secured for ease of management.

1. Go to your Tableau instance, click on your user profile on the top right, then click on *`My Account Settings`*.
2. There, scroll down until you find *`Personal Access Tokens`.*
3. In `Token Name`, type a suitable name, we recommend `decube data observability`. Then click on *`Create New Token`.*
4. A popup will be displayed where you can view `Token Name` and `Token Secret`. Click on *`Copy to clipboard`*. We recommend saving this credential, like on a password manager as the `Token Secret` will only be shown to you once.

## Additional configuration for lineage

To build out the lineage, we will need to know the data sources that you've referenced within your Tableau definitions. You can do a one-time mapping for the sources by using the Additional Config in the Modify Data Sources page. Check out how below.

1. Go to My Account > Data Sources.

<figure><img src="/files/mrs6Xpksa97rFo0N4DwC" alt=""><figcaption></figcaption></figure>

2. Click on Modify.

<figure><img src="/files/CAIKet81XFGRMwFh0vaF" alt=""><figcaption></figcaption></figure>

3. Click on the Dropdown for Select Data Sources. Select the data sources (4) that your Tableau objects are referencing to and click on `Save preferences` (5).

<figure><img src="/files/nURKzne0iapdAUEOo9ea" alt=""><figcaption></figcaption></figure>

### FAQ

#### I am not able to see the lineage from my data source to Tableau.

1. Check in the Tableau Settings in the General tab that `Automatic Access to Metadata about Databases and Tables` is enabled in your organization.

<figure><img src="/files/37aBIjI8UbKJLAZ1W2pz" alt=""><figcaption></figcaption></figure>

2. If it is enabled, it may be that the access for the user used to generate the Personal Access Token in the step above needs to have its access elevated in the Project. You will need to grant the `Project Leader` access to the user. Read more about the permissions that the `Project Leader` has in this [Tableadocumentation](https://help.tableau.com/current/server/en-us/permissions_projects.htm).


# Looker

Add a Looker connection to your decube so that you can discover the lineage of your assets from source to dashboards.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Chart, Dashboard
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure
  {% endtab %}
  {% endtabs %}

Looker supports lineage mapping from the following sources:

* Upstream Connectors: mysql, singlestore, postgresql, redshift, snowflake, azure\_server, sql\_server, synapse, bigquery, databricks, oracle

## Connection Requirements

## Minimum Requirement

From our Looker account, we will need the following information:

* Looker instance URL
* Client ID
* Client secret

<figure><img src="/files/pYPQpLeU4uoHGz9Q3fc0" alt=""><figcaption><p>Looker</p></figcaption></figure>

To ensure proper functionality of the ingestion process, it is imperative to grant the following permissions.

* access\_data
* develop
* explore
* manage\_project\_models
* manage\_project\_connections
* see\_lookml
* see\_lookml\_dashboards
* see\_looks
* see\_queries
* see\_sql

For more info on how to create permission sets. Please refer to this looker [documentation](https://cloud.google.com/looker/docs/admin-panel-users-roles#creating_permission_sets).

## Additional configuration for lineage

To build out the lineage, we will need to know the data sources that you've referenced within your Looker definitions. You can do a one-time mapping for the sources by using the Additional Configuration in the Data Sources page. Check out how below.

{% content-ref url="/pages/NtHz59Y2JYp680p0fWX7" %}
[Additional configurations](/transformation-tools/additional-configurations)
{% endcontent-ref %}

For Looker, we match your Views in Looker to decube Tables by the **View name**. If you have customized View names, please refer to the next section below.

### Override View Mapping (Optional)

<figure><img src="/files/inMGYgB16EJKiOx69VNT" alt=""><figcaption><p>Mapping for Looker Lineage which is optional</p></figcaption></figure>

If you have a customized View name, it is recommended to use the View Mapping Override to add the name of the Table in decube so that our metadata scanner can build the lineage automatically to your Looker view.


# PowerBI

Add a PowerBI connection to your decube so that you can discover the lineage of your assets from source to dashboards.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Virtual Table, Virtual Column, Chart, Dashboard
* **Configurable Collection** — selective ingestion of schemas/workspaces in Data Source Management
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Preview
* Data Quality
* External Table
* View Table
* Stored Procedure
  {% endtab %}
  {% endtabs %}

Connecting to PowerBI requires credentials that can be configured and found through this guide. These credentials include:

* `Tenant ID`
* `Client ID`
* `Client Secret`

<figure><img src="/files/COa1nYGwqX0jT5JMhXYd" alt=""><figcaption><p>PowerBI</p></figcaption></figure>

To note, for complete Lineage to report sources, with [additional configurations](/transformation-tools/additional-configurations) we only support sources from:

* BigQuery
* PostgreSQL
* Microsoft SQL
* Snowflake
* Synapse

{% hint style="warning" %}
**The connector might be rate-limited when it is scanning more than 1000 workspaces.**
{% endhint %}

### Prerequisite

1. Access to Azure Active Directory for service principals.
2. A PowerBI Pro or Premium (Premium per-user or Premium capacity ) workspace.
3. Admin access to PowerBI to change Tenant Settings.

### Creating Service Principals and Security Group

1. Go to Azure Active Directory, under the Manage tab to the left of the screen, look for `App Registrations` and click it.
2. Click on `New Registrations` on the top of the tab and you'll be presented with the screen below.

<figure><img src="/files/5ksMPRfqxxOEPJm3uU16" alt=""><figcaption></figcaption></figure>

3. We recommend entering `decube` as the Name. Click Register when you are done.

<figure><img src="/files/KlD9gjvh5N0avsm5qGhJ" alt=""><figcaption></figcaption></figure>

4. Click Overview on the left tab and take note of the `Client ID` and `Tenant ID` we would require when setting up decube.
5. Click on `Certificates and Secret`s on the left tab again and click on `New Client Secret` as below. We recommend the description to be `'decube client secret`' and the expiry date to be until the end of your contract with us.

**Make sure to copy the Value of the Client Secret and store it somewhere safe like Azure Key Vault. We will need this value for registration.**

<figure><img src="/files/kPfOb9KwhVcr8w5vA3J3" alt=""><figcaption></figcaption></figure>

6. Go back to the Azure Active Directory main page and on the left tab again click on Groups. Click `Create New Group`.

<figure><img src="/files/UiCSv4RXgVeF9J3tYa4Q" alt=""><figcaption></figcaption></figure>

7. Ensure Group Type is `Security.`
8. Click on members and search for the previously configured Service Principal. Here it would be `decube`.

<figure><img src="/files/QdcZaBoNDo2JxJ2VzXzv" alt=""><figcaption></figcaption></figure>

9. Click on `Select` to add the service principal into the group.
10. Click on `Create` and we are done retrieving the `Tenant ID`, `Client ID` and `Client Secret` for this integration.

### Setting up and providing permissions to Service Principal

1. On the main PowerBI workspace page, on the top right corner, click on the gear icon to Setting and below `Governance and Insight` you will find `Admin Portal`.
2. Here in the Admin Portal, find Tenant Settings and search for `Allow service principals to use Power BI APIs under Developer Settings`.
3. Enable it and either `Apply to Entire Organisation` or `Specific Security Group.` Here you can search for the one we created earlier. Example of how it should look like below.

<figure><img src="/files/TVSbIGcSFWxNe5uyyDux" alt=""><figcaption></figcaption></figure>

4. Please enable and do this for the following settings as well under `Admin API` settings.

* `Allow service principals to use read-only admin APIs`
* `Enhance admin APIs responses with detailed metadata`
* `Enhance admin APIs responses with DAX and mashup expressions`

### Additional Configuration for lineage

To build out the lineage, we will need to know the data sources that you've referenced within your PowerBI definitions. You can do a one-time mapping for the sources by going to the Data Sources page > Modify Data Source > Additional Config.

<figure><img src="/files/PXogfIA2n9FlCO2UJWIs" alt=""><figcaption></figcaption></figure>


# AWS S3

Connect your S3 to see your S3 datasets and files within the Catalog.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column
* **Preview** — sample data preview
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure
  {% endtab %}
  {% endtabs %}

To connect your AWS S3 to Decube, we will need the following information.

Choose authentication method:

a. [**AWS Identity**](#a.-aws-roles):

* Select AWS Identity
* Customer AWS Role ARN
* Region
* Path Specs
* Data source name

<figure><img src="/files/8wP4BwoROe1jdmb2uYsI" alt=""><figcaption><p>S3 Datalake using AWS Identity<br></p></figcaption></figure>

<figure><img src="/files/dDbj2P2eOwpAXyXi24bH" alt=""><figcaption><p>Continuation of S3 Data Lake setup using AWS Identity</p></figcaption></figure>

b. **AWS** **Access** **Key**:

* Access Key ID
* Secret Access Key
* Region
* Path Specs
* Data source name

<figure><img src="/files/qzrABC68xoYdMHWJ4SqZ" alt=""><figcaption><p>S3 Datalake using AWS Access Key</p></figcaption></figure>

## Connection Options:

#### a. AWS Roles

{% hint style="info" %}
This section will create a **Customer AWS Role** within your AWS account that has the right set of permission to access your data sources.
{% endhint %}

* Step 1: Go to your AWS Account > IAM Module > Roles
* Step 2: Click on **Create role**

<figure><img src="/files/oqh7ru346tgegU5ag3yO" alt=""><figcaption></figcaption></figure>

* Step 3: Choose **Custom trust policy**

<figure><img src="/files/69oNLlXEpHGoPBgwyZUV" alt=""><figcaption></figcaption></figure>

* Step 4: Specify the following as the trust policy, replacing `DECUBE-AWS-IDENTITY-ARN` and `EXTERNAL-ID` with values from [AWS Identities](/security-and-connectivity/aws-identities#generating-a-decube-aws-identity).

```
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "AWS": "<DECUBE-AWS-IDENTITY-ARN>"
            },
            "Action": "sts:AssumeRole",
            "Condition": {
                "StringEquals": {
                    "sts:ExternalId": "<EXTERNAL-ID>"
                }
            }
        }
    ]
}
```

* Step 5: Click next to proceed to attach policy.
* Step 6: Click on Attach Policies and Create Policy and choose JSON Editor. Input the following policy and press next, input the policy name of your choice and press Create Policy.

```
{
	"Version": "2012-10-17",
	"Statement": [
		{
			"Sid": "VisualEditor0",
			"Effect": "Allow",
			"Action": [
				"s3:GetObject",
				"s3:ListBucket",
				"s3:ListAllMyBuckets"
			],
			"Resource": [
				"arn:aws:s3:::{bucket-name}",
				"arn:aws:s3:::{bucket-name}/*"
			]
		}
	]
}
```

#### b. AWS IAM User

* Step 1: Login to AWS Console and proceed to IAM > User > Create User

<figure><img src="/files/Z4zzj2GNJr59njp5xexw" alt=""><figcaption></figcaption></figure>

* Step 2: Click on Attach Policies and Create Policy and choose JSON Editor input the following policy and press next, input the policy name of your choice and press Create Policy

```jsx
{
	"Version": "2012-10-17",
	"Statement": [
		{
			"Sid": "VisualEditor0",
			"Effect": "Allow",
			"Action": [
				"s3:GetObject",
				"s3:ListBucket",
				"s3:ListAllMyBuckets"
			],
			"Resource": [
				"arn:aws:s3:::{bucket-name}",
				"arn:aws:s3:::{bucket-name}/*"
			]
		}
	]
}
```

* Step 3: Search for the policy you created just now, select it and press Next.

<figure><img src="/files/6bFP9p4QlYOBZ3tqvvXN" alt=""><figcaption></figcaption></figure>

* Step 4: Press **Create user**

<figure><img src="/files/ZeOEu22NVy8UcDftq1To" alt=""><figcaption></figcaption></figure>

* Step 5: Navigate to the newly created user and click on `Create access key`

<figure><img src="/files/nOxqAlhzb2DFdD3aa6Jy" alt=""><figcaption></figcaption></figure>

* Step 6: Choose `Application running outside AWS`

<figure><img src="/files/zMbKBrIXIDuwOdNde7kS" alt=""><figcaption></figcaption></figure>

* Step 7: Save the provided access key and secret access key. You will not be able to retrieve these keys again

<figure><img src="/files/7LAgiaHpJAyGV1WyxMdT" alt=""><figcaption></figcaption></figure>

#### AWS KMS

If the bucket intended to be connected to Decube is encrypted using a customer managed KMS key, you will need to add the AWS IAM user created above to the key policy statement.

1. Login to AWS Console and proceed to AWS KMS > Customer-managed keys.
2. Find the key that was used to encrypt the AWS S3 bucket.
3. On the Key policy tab, click on `Edit`

<figure><img src="/files/xFbswZjG4fJ8PzhQS723" alt=""><figcaption></figcaption></figure>

4. Assuming the user created is `decube-s3-datalake`

a. If there is not an existing policy attached to the key

```bash
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "Allow decube to use key",
            "Effect": "Allow",
            "Principal": {
                "AWS": [
                    "arn:aws:iam::<AWSAccountID>:user/{decube-s3-datalake}"
                ]
            },
            "Action": "kms:Decrypt",
            "Resource": "*"
        }
    ]
}

```

b. If there is an existing policy, append this section to the `Statement` array:

```bash
{
    "Statement": [
        {
            "Sid": "Allow decube to use key",
            "Effect": "Allow",
            "Principal": {
                "AWS": [
                    "arn:aws:iam::<AWSAccountID>:user/{decube-s3-datalake}"
                ]
            },
            "Action": "kms:Decrypt",
            "Resource": "*"
        }
    ]
}
```

5. `Save Changes`

## Path Specs

Path Specs (`path_specs`) is a list of Path Spec (`path_spec`) objects where each individual `path_spec` represents one or more datasets. Providing a path specification represents the formatted path to the dataset in your S3 bucket which Decube will use to ingest and catalog the data.

The provided path specification MUST end with `*.*` or `*.[ext]` to represent leaf level. (Note that here '\*' is *not* a wildcard symbol). If `*.[ext]` is provided then files with only specified extension type will be scanned. "`.[ext]`" can be any of the supported file types listed below.

Each `path_spec` represents only one file type (e.g., only `*.csv` or only `*.parquet`). To ingest multiple file types, add multiple `path_spec` entries.

**SingleFile pathspec**: PathSpec without `{table}` (targets individual files).**MultiFile pathspec**: PathSpec with `{table}` (targets folders as datasets).

**Include only datasets that match this pattern**: If the path spec specifies a table, and the regex is provided, only datasets that match the regex will be included.

### File Format Settings

* **CSV**
  * `delimiter` (default: `,`)
  * `escape_char` (default: `\\`)
  * `quote_char` (default: `"`)
  * `has_headers` (default: true)
  * `skip_n_line` (default: 0)
  * `file_encoding` (default: UTF-8; supported: ASCII, UTF-8, UTF-16, UTF-32, Big5, GB2312, EUC-TW, HZ-GB-2312, ISO-2022-CN, EUC-JP, SHIFT\_JIS, CP932, ISO-2022-JP, EUC-KR, ISO-2022-KR, Johab, KOI8-R, MacCyrillic, IBM855, IBM866, ISO-8859-5, windows-1251, MacRoman, ISO-8859-7, windows-1253, ISO-8859-8, windows-1255, TIS-620)
* **Parquet**
  * No options
* **JSON/JSONL**
  * `file_encoding` (default: UTF-8; see above for supported encodings)
* **Delta Table**
  * When selecting Format = `Delta Table`, the path spec MUST include the named token `{table}`. The connector expects the `{table}` token in the path spec so it can discover Delta table roots. No per-file options are required.

**Additional points to note**

* Folder names should not contain {, }, \*, / in their names.
* Named variable {folder} is reserved for internal working. please do not use in named variables.

### Example Path Specs

#### Example 1 - Individual file as Dataset (SingleFile pathspec)

Bucket structure:

```
test-bucket
├── employees.csv
├── departments.json
└── food_items.csv
```

Path specs config to ingest `employees.csv` and `food_items.csv` as datasets:

```
path_specs:
    - s3://test-bucket/*.csv
```

This will automatically ignore `departments.json` file. To include it, use `*.*` instead of `*.csv`.

**Example 2 - Folder of files as Dataset (without Partitions)**

Bucket structure:

```
test-bucket
└──  offers
     ├── 1.csv
     └── 2.csv
```

Path specs config to ingest folder `offers` as dataset:

```
path_specs:
    - s3://test-bucket/{table}/*.csv
```

`{table}` represents folder for which dataset will be created.

**Example 3 - Folder of files as Dataset (with Partitions)**

Bucket structure:

```
test-bucket
├── orders
│   └── year=2022
│       └── month=2
│           ├── 1.parquet
│           └── 2.parquet
└── returns
    └── year=2021
        └── month=2
            └── 1.parquet

```

Path specs config to ingest folders `orders` and `returns` as datasets:

```
path_specs:
    - s3://test-bucket/{table}/*/*/*.parquet
```

**Example 4 - Advanced - Either Individual file OR Folder of files as Dataset**

Bucket structure:

```
test-bucket
├── customers
│   ├── part1.json
│   ├── part2.json
│   ├── part3.json
│   └── part4.json
├── employees.csv
├── food_items.csv
├── tmp_10101000.csv
└──  orders
    └── year=2022
        └── month=2
            ├── 1.parquet
            ├── 2.parquet
            └── 3.parquet

```

Path specs config:

```
path_specs:
    - path_spec_1: s3://test-bucket/*.csv
    - path_spec_2: s3://test-bucket/{table}/*.json
    - path_spec_3: s3://test-bucket/{table}/*/*/*.parquet

```

Above config has 3 path\_specs and will ingest following datasets

* `employees.csv` - Single File as Dataset
* `food_items.csv` - Single File as Dataset
* `customers` - Folder as Dataset
* `orders` - Folder as Dataset

**Valid path\_specs.include**

```python
s3://my-bucket/foo/tests/bar.csv # single file table
s3://my-bucket/foo/tests/*.* # mulitple file level tables
s3://my-bucket/foo/tests/{table}/*.parquet #table without partition
s3://my-bucket/foo/tests/{table}/*/*.csv #table where partitions are not specified
s3://my-bucket/foo/tests/{table}/*.* # table where no partitions as well as data type specified
s3://my-bucket/{dept}/tests/{table}/*.parquet # specifying keywords to be used in display namepartition key and value format
```

#### Example 5 - Delta Table (S3)

For Delta Table support, include `{table}` in the path. Example simple path spec for S3:

```
path_specs:
    - s3://my-bucket/{table}/
```

The connector will interpret `{table}` as the table root for each Delta table.

#### Supported file types

* CSV (`*.csv`)
* JSON (`*.json`)
* JSONL (`*.jsonl`)
* Parquet (`*.parquet`)
* Delta (`Delta Table`)

### Notes

* Data Quality Monitoring is no longer supported for S3 sources. Only cataloging is available.
* For advanced dataset structures, add multiple `path_spec` entries as needed, each with its own file type and settings.


# Azure Data Lake Storage (ADLS)

Azure Data Lake Storage (ADLS) is a scalable and secure data lake solution from Microsoft Azure designed to handle the vast amounts of data generated by modern applications.

## Supported Capabilities

{% tabs %}
{% tab title="Supported Capabilities" %}
**General**

* **Metadata** — metadata extraction and display of asset information (tables, columns, schemas). Types collected: Schema, Table, Column
* **Preview** — sample data preview
  {% endtab %}

{% tab title="Not Supported" %}
**General**

* Profiling
* Data Quality
* Configurable Collection
* External Table
* View Table
* Stored Procedure
  {% endtab %}
  {% endtabs %}

To connect your ADLS to Decube, the following information is required:

* Tenant ID
* Client ID
* Client Secret

{% hint style="info" %}
**Potential Data Egress**

Under the SaaS deployment model, data must be transferred from the storage container to the Data Plane to inspect files, retrieve schema information, and perform data quality monitoring.\
\
If this is not preferable, you may opt for a Self-Hosted deployment model or bring your own Azure Function [Azure Function for Metadata](/datalake/azure-data-lake-storage-adls/azure-function-for-metadata)
{% endhint %}

## **Firewall and connectivity configuration**

By default, Azure Storage accounts may allow access from all networks. However, if your organization requires Public network access to be disabled for security compliance, you must explicitly whitelist Decube's IP addresses to allow our connectors to access your Data Lake.

Follow the steps below to configure your firewall settings.

**1. Navigate to Networking Settings**

1. Log in to the Azure Portal and navigate to your Storage Account.
2. In the left-hand sidebar, under Security + networking, select Networking.
3. Under the Firewalls and virtual networks tab, locate the Public network access setting.

**2. Enable Access for Selected Networks**

To allow Decube to connect while keeping the storage account private from the general public:

1. Select Enabled from selected virtual networks and IP addresses.
2. This option enables the Firewall section below it, where you can specify allowed IP addresses.

<figure><img src="/files/7GwxUDkyVrVT2zjnxt3I" alt=""><figcaption></figcaption></figure>

**3. Whitelist Decube IP Addresses**

In the Firewall section, you must add the IP addresses corresponding to the region where your Decube SaaS instance is hosted. See the section on IP Whitelisting to get the list of IP address.

{% content-ref url="/pages/IeabSxBfHBW69C0lrCKD" %}
[IP Whitelisting](/security-and-connectivity/ip-whitelisting)
{% endcontent-ref %}

## **Credentials setup**

### Setup on Microsoft Azure

1. On the Azure Home Page, go to `Azure Active Directory`. The **Tenant ID** can be copied from the Basic information section.\\

   <figure><img src="/files/ABScEep7hx5OlqithScq" alt=""><figcaption></figcaption></figure>
2. Go to `App registrations`.\\

   <figure><img src="/files/Qpp7fnWb5B6ZgxN6Yfma" alt=""><figcaption></figcaption></figure>
3. Click on `New registration`.\\

   <figure><img src="/files/Ng9XTy6MhieZJS5HlYWK" alt=""><figcaption></figcaption></figure>
4. Click `Register.`

   <figure><img src="/files/1446SIv6UEGeFwzt8wfs" alt=""><figcaption></figcaption></figure>
5. Save the `Application (client) ID` and `Directory (tenant) ID`.
6. Click `Add a certificate or secret`.
7. Go to `Client secrets` and client `+ New client secret`.\\

   <figure><img src="/files/kQuP5uipdF2iuRKmcOTu" alt=""><figcaption></figcaption></figure>
8. Click `+Add`.\
   \\

   <figure><img src="/files/Gr32F5bBLkl9ILLTWjeO" alt=""><figcaption></figcaption></figure>
9. Copy and save the `Value` for the **client secret**.\\

   <figure><img src="/files/ywHel0a6wMQBd9mVKKOB" alt=""><figcaption></figcaption></figure>

#### **Assigning Role to Credentials**

1. From Azure Services, find and click on Storage Accounts. You should be able to see the option for Access control (IAM) on the left sidebar.

<figure><img src="/files/dYPJr9pySNVnQJO8pMfc" alt=""><figcaption></figcaption></figure>

2. Click on Access Control -> Click on '+Add' -> Click on Role assignments.

<figure><img src="/files/98ov6VjEOTjONcUVNdBZ" alt=""><figcaption></figcaption></figure>

3. Find the role called **Storage Blob Data Reader** click on it and click next.
4. On the next page, search for the name of the application that you just created on Microsoft Entra ID.
5. Assign it to the role.

<figure><img src="/files/yi5peBUNbsZ5ETogPzRL" alt=""><figcaption></figcaption></figure>

## Path Specs

Path Specs (`path_specs`) is a list of Path Spec (`path_spec`) objects where each individual `path_spec` represents one or more datasets for cataloging in ADLS.

The provided path specification MUST end with `*.*` or `*.[ext]` to represent the leaf level. (Note: here `*` is not a wildcard symbol.) If `*.[ext]` is provided, only files with the specified extension will be scanned. `.[ext]` can be any of the supported file types listed below.

Each `path_spec` represents only one file type (e.g., only `*.csv` or only `*.parquet`). To ingest multiple file types, add multiple `path_spec` entries.

**SingleFile pathspec**: PathSpec without `{table}` (targets individual files).**MultiFile pathspec**: PathSpec with `{table}` (targets folders as datasets).

### PathSpec Structure

* Take note of thesee following parameters when building a path spec:
  * Storage account name
  * Container name
  * Folder path

<figure><img src="/files/BYxcImVuBSn9nXz6pTL8" alt=""><figcaption></figcaption></figure>

Follow this schema when building a path spec:

```bash
"abfs://{container name}@{storage account name}.dfs.core.windows.net/{folder path}"
example
"abfs://first@decubetestadls.dfs.core.windows.net/second/*.*"// Some code
```

**Include only datasets that match this pattern**: If the path spec specifies a table and a regex is provided, only datasets that match the regex will be included.

### File Format Settings

* **CSV**
  * `delimiter` (default: `,`)
  * `escape_char` (default: `\\`)
  * `quote_char` (default: `"`)
  * `has_headers` (default: true)
  * `skip_n_line` (default: 0)
  * `file_encoding` (default: UTF-8; supported: ASCII, UTF-8, UTF-16, UTF-32, Big5, GB2312, EUC-TW, HZ-GB-2312, ISO-2022-CN, EUC-JP, SHIFT\_JIS, CP932, ISO-2022-JP, EUC-KR, ISO-2022-KR, Johab, KOI8-R, MacCyrillic, IBM855, IBM866, ISO-8859-5, windows-1251, MacRoman, ISO-8859-7, windows-1253, ISO-8859-8, windows-1255, TIS-620)
* **Parquet**
  * No options
* **JSON/JSONL**
  * `file_encoding` (default: UTF-8; see above for supported encodings)
* **Delta Table**
  * When selecting Format = `Delta Table`, the path spec MUST include the named token `{table}`. The connector expects the `{table}` token in the path spec so it can discover Delta table roots. No per-file options are required.

**Additional points to note**

* Folder names should not contain {, }, \*, / in their names.
* Named variable {folder} is reserved for internal working. Please do not use in named variables.

### Example Path Specs

#### Example 1 - Individual file as Dataset (SingleFile pathspec)

Bucket structure:

```
test-bucket
├── employees.csv
├── departments.json
└── food_items.csv
```

Path specs config to ingest `employees.csv` and `food_items.csv` as datasets:

```
path_specs:
    - include: abfs://test-container@test-storage-account.dfs.core.windows.net/*.csv
```

This will automatically ignore `departments.json` file. To include it, use `*.*` instead of `*.csv`.

**Example 2 - Folder of files as Dataset (without Partitions)**

Bucket structure:

```
test-bucket
└──  offers
     ├── 1.csv
     └── 2.csv
```

Path specs config to ingest folder `offers` as dataset:

```
path_specs:
    - include: abfs://test-container@test-storage-account.dfs.core.windows.net/{table}/*.csv
```

`{table}` represents folder for which dataset will be created.

**Example 3 - Folder of files as Dataset (with Partitions)**

Bucket structure:

```
test-bucket
├── orders
│   └── year=2022
│       └── month=2
│           ├── 1.parquet
│           └── 2.parquet
└── returns
    └── year=2021
        └── month=2
            └── 1.parquet

```

Path specs config to ingest folders `orders` and `returns` as datasets:

```
path_specs:
    - include: abfs://test-container@test-storage-account.dfs.core.windows.net/{table}/*/*/*.parquet
```

**Example 4 - Advanced - Either Individual file OR Folder of files as Dataset**

Bucket structure:

```
test-bucket
├── customers
│   ├── part1.json
│   ├── part2.json
│   ├── part3.json
│   └── part4.json
├── employees.csv
├── food_items.csv
├── tmp_10101000.csv
└──  orders
    └── year=2022
        └── month=2
            ├── 1.parquet
            ├── 2.parquet
            └── 3.parquet

```

Path specs config:

```
path_specs:
    - path_spec_1: abfs://test-container@test-storage-account.dfs.core.windows.net/*.csv
    - path_spec_2: abfs://test-container@test-storage-account.dfs.core.windows.net/{table}/*.json
    - path_spec_3: abfs://test-container@test-storage-account.dfs.core.windows.net/{table}/*/*/*.parquet

```

Above config has 3 path\_specs and will ingest following datasets

* `employees.csv` - Single File as Dataset
* `food_items.csv` - Single File as Dataset
* `customers` - Folder as Dataset
* `orders` - Folder as Dataset and will ignore file `tmp_10101000.csv`

**Valid path\_specs.include**

<pre class="language-python"><code class="lang-python"><strong>abfs://test-container@test-storage-account.dfs.core.windows.net/foo/tests/bar.csv # single file table
</strong>abfs://test-container@test-storage-account.dfs.core.windows.net/foo/tests/*.* # mulitple file level tables
abfs://test-container@test-storage-account.dfs.core.windows.net/foo/tests/{table}/*.parquet #table without partition
abfs://test-container@test-storage-account.dfs.core.windows.net/tests/{table}/*/*.csv #table where partitions are not specified
abfs://test-container@test-storage-account.dfs.core.windows.net/tests/{table}/*.* # table where no partitions as well as data type specified
abfs://test-container@test-storage-account.dfs.core.windows.net/{dept}/tests/{table}/*.parquet # specifying keywords to be used in display namepartition key and value format
</code></pre>

#### Example 5 - Delta Table (ADLS)

For Delta Table support, include `{table}` in the path. Example simple path spec for ADLS:

```
path_specs:
    - abfss://datalake@account.dfs.core.windows.net/{table}/
```

The connector will interpret `{table}` as the table root for each Delta table.

#### Supported file types

* CSV (`*.csv`)
* JSON (`*.json`)
* JSONL (`*.jsonl`)
* Parquet (`*.parquet`)
* Delta (`Delta Table`)

### Notes

* Data Quality Monitoring is no longer supported for ADLS sources. Only cataloging is available.
* For advanced dataset structures, add multiple `path_spec` entries as needed, each with its own file type and settings.


# Azure Function for Metadata

Decube can leverage a customer-hosted Azure Function to retrieve file schemas, minimizing file egress from the customer's environment to Decube's compute. While this approach is generally applicable to the SaaS deployment model, it can also be beneficial for self-hosted solutions where the Data Plane resides in a different region than the Storage Container.

## Prerequisite

* Create an Azure Function with Custom Handler as the Runtime Stack
* Networking between Decube Data Plane (either SaaS or customer own) to the Azure Function

## Creating Azure Function

1. Go to your Azure Portal and navigate to Function App and click on \`Create\`

<figure><img src="/files/iSXA3tdEKTvS3pzzTZt6" alt=""><figcaption></figcaption></figure>

2. Choose either Consumption (recommended), Functions Premium or App Service

<figure><img src="/files/3jIGptC9sqEbDbkMn2QF" alt=""><figcaption></figcaption></figure>

3. Fill in the appropriate field
   1. Function App Name - a unique name for this function
   2. Runtime stack - Custom Handler
   3. Region - Closest to your ADLS storage account region
   4. Operating System - Linux

<figure><img src="/files/hoQIgZWERFvY4sqVsdox" alt=""><figcaption></figcaption></figure>

4. Under Networking, choose Enable Public Access - On
   1. See [Azure documentation](https://learn.microsoft.com/en-us/azure/azure-functions/functions-networking-options?tabs=azure-portal) if alternate networking is required

<figure><img src="/files/Zmbpk3jLP0VJTBV1RSFA" alt=""><figcaption></figcaption></figure>

5. Under Deployment, disable Continous Deployment
6. Storage, Monitoring, Tags are to be set up according to customer need
7. Create the Azure Function

## Deployment

Visit <https://github.com/DecubeIO/adls-azure-function> to start

1. Clone the repo to a local machine
2. Run `func azure functionapp publish $functionAppName --custom`
   1. `$functionAppName`is based on the Azure Function name used in [#creating-azure-function](#creating-azure-function "mention")

## Using the Azure Function

<figure><img src="/files/Ezr9PajBI3ZdZyryxzfM" alt=""><figcaption></figcaption></figure>

1. When creating/modifying an ADLS source, enable `Use remote Azure Function`
2. Fill in the Azure Function URL in this format: `https://myfunction.azurewebsites.net`
3. Fill in the Azure Function Key. This can be found in your Portal Azure > Go to created function above > Functions > App Keys > Either choose an existing host key or create a new host key


# Google Cloud Storage (GCS)

{% hint style="info" %}
This connector is coming soon.
{% endhint %}


# Overview

The Universal Catalog Framework let you document data sources without native Decube connectors, so they participate in your catalog, lineage, and governance workflows.

Decube supports two models for bringing source metadata into your catalog: automated ingestion via a native connector, and manual ingestion via virtual sources. Virtual sources give you full control — you push the metadata to Decube directly through the UI or API, on your own schedule, without needing a connector. This makes virtual sources the right choice for tightly secured systems, on-premises infrastructure, or any source where you prefer to own and manage exactly what metadata enters Decube.

{% embed url="<https://youtu.be/OJg_VybvQ6c?si=4RTbYMZ59Ilw3mWY>" %}

## What are virtual sources?

A virtual source is a data source whose metadata you manage directly in Decube, rather than having Decube pull it from the system automatically. You populate objects using Decube's existing data model through the UI or API, giving you precise control over what is documented and how. Virtual sources and the objects under them are fully searchable in the Catalog, can be linked to Glossary terms, and support manual lineage edges to any other asset in your workspace.

## Supported object types

Each virtual source can contain the following types of objects:

| Hierarchy                        | Object types                          |
| -------------------------------- | ------------------------------------- |
| Source → Schema → Table → Column | Schema, Virtual Table, Virtual Column |
| Source → Data Job → Data Task    | Data Job, Data Task                   |
| Source → Dashboard               | Dashboard                             |
| Source → Chart                   | Chart                                 |

## What's not available for virtual objects

Virtual sources don't support operations that require a live connection to the underlying system:

* **Profiling and sampling** — no Preview or Profile tabs on virtual assets
* **Data Quality monitors** — Monitors and Incidents tabs are unavailable on virtual assets
* **Automated lineage** — lineage must be added manually

All other catalog features work as they do on native assets: documentation, custom attributes, ownership, manual lineage, and Glossary links.

## Permissions

Creating, editing, and deleting virtual sources and their objects requires the **Manage data sources** permission (`admin_data_source:edit`). Viewing virtual sources requires view access to the source. If you have view but not edit access, all editing controls are visible but disabled in the UI.

***

{% content-ref url="/pages/HgudG62v5zAZqLrMERQK" %}
[Manage Virtual Sources](/virtual-sources-and-objects/manage-virtual-sources)
{% endcontent-ref %}

{% content-ref url="/pages/g96V6dpECdQFUkW5mkAH" %}
[Manage Virtual Objects](/virtual-sources-and-objects/manage-virtual-objects)
{% endcontent-ref %}

{% content-ref url="/pages/AOBuHq3Rf5WZIUSynFE6" %}
[Virtual Sources in Lineage](/virtual-sources-and-objects/virtual-sources-in-lineage)
{% endcontent-ref %}


# Manage Virtual Sources

Create, edit, enable or disable, and delete virtual sources from the Integrations page.

Virtual sources live on the **Integrations** page under the **Data Sources** tab, alongside your native connector sources. You can create and manage them entirely from this page.

## Create a virtual source

1. Navigate to **Integrations** and select the **Data Sources** tab.
2. Click the **Create new** button and select **Create a virtual source** from the dropdown.

<figure><img src="/files/VFR2RrCL3ebn2xEEPVVm" alt=""><figcaption></figcaption></figure>

3. Enter a **Name** for the source. The name must be unique across all sources in your workspace.
4. Click **Create** to finish, or click **Add virtual assets** to go directly to the Asset Tree Editor and start building your object hierarchy.

<figure><img src="/files/zChQjdxqpCGuXOqU5HKS" alt=""><figcaption></figcaption></figure>

After creation, the virtual source appears in the Data Sources list with a virtual source icon. The **Children updated at** timestamp on the row updates each time any object under the source is created, edited, or deleted.

### Virtual source landing page

After creation you land on the virtual source page. From here you can start populating the source with objects via:

* **UI** — opens the Asset Tree Editor to build the object hierarchy interactively. See [Manage Virtual Objects](/virtual-sources-and-objects/manage-virtual-objects).
* **API** — links to the API reference; the virtual source ID is shown here for use in API calls. See [Virtual Sources](/public-api/overview/index/virtual-sources).

You can also connect virtual objects into your lineage graph from their asset detail pages. See [Virtual Sources in Lineage](/virtual-sources-and-objects/virtual-sources-in-lineage).

## Edit a virtual source

1. On the **Data Sources** tab, locate the virtual source row and click **Manage**.
2. In the modal, update any of the following fields:
   * **Name**
   * **Description**
   * **Source owner** — defaults to the creator; can be transferred to any other user
   * **Enabled / Disabled** state

Changes take effect immediately on save.

## Enable or disable a virtual source

Disabling a virtual source removes it from Catalog search results and lineage diagrams without deleting it. This behaves the same as enabling and disabling native sources.

Open the **Manage** modal for the source and toggle the enabled state on or off.

{% hint style="warning" %}
Disabling a virtual source hides it and all its objects from the Catalog and lineage graphs. Re-enabling it restores full visibility.
{% endhint %}

## Delete a virtual source

Deleting a virtual source is a **hard delete** — the source and all virtual objects under it are permanently removed from the database.

{% hint style="danger" %}
This action is irreversible. All virtual objects under the source, including their metadata, are permanently deleted.
{% endhint %}

1. Open the **Manage** modal for the virtual source.
2. Click **Delete source** and confirm in the secondary confirmation step.

{% hint style="info" %}
Virtual object deletion works differently — deleting an individual object is a soft delete that removes it from the UI but retains the record in the database. See [Manage Virtual Objects](/virtual-sources-and-objects/manage-virtual-objects) for details.
{% endhint %}


# Manage Virtual Objects

Add, edit, and delete virtual objects under a virtual source using the Asset Tree Editor.

{% hint style="info" %}
This article describes the method to update virtual objects via the UI. If you intend to develop a workflow to update via API instead, check out our [Public API documentation](/public-api/overview/index/virtual-sources).
{% endhint %}

Virtual objects are the assets that live under a virtual source — schemas, tables, columns, data jobs, data tasks, dashboards, and charts. You create and manage them using the **Asset Tree Editor**, which auto-saves every change as you make it.

Open the Asset Tree Editor from the **virtual source landing page** or from any **virtual asset's detail page**. The editor opens scoped to wherever you launched it from — from the virtual source you see all object types across four tabs (Schemas, Data Jobs, Dashboards, Charts); from an individual asset you see that asset and its children.

<figure><img src="/files/oiXmXPXsXoLjDEVxwedd" alt=""><figcaption></figcaption></figure>

## Add an object

Each tab has an **Add virtual \[type]** button that creates a new row at the top of the list. Enter a name and click away to save. To add a child object (e.g. a table under a schema), expand the parent row first — the **Add child** button appears inside it.

## Edit an object

Click any field in a row to edit it. Changes save automatically when you click away. For the **Definition** field, clicking it opens a markdown editor modal where you write and save the content.

## Delete an object

Click the delete icon on any row and confirm in the modal. Deleting a parent object (e.g. a Schema) also removes all its children.

{% hint style="warning" %}
Deletion is irreversible — the object and all its metadata are removed from the Catalog, lineage, and search results.
{% endhint %}

## Object type reference

| Asset type | Fields                                                                                                                                              |
| ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
| Schema     | `name`                                                                                                                                              |
| Table      | `name`, `definition` (markdown, renders in the Definition tab)                                                                                      |
| Column     | `name`, `unified_type` (data type), `raw_type` (free text, e.g. `VARCHAR(255)`), `constraints` (`NOT_NULL`, `UNIQUE`, `PRIMARY_KEY`, `FOREIGN_KEY`) |
| Data Job   | `name`, `schedule`, `definition` (markdown, renders in the Definition tab)                                                                          |
| Data Task  | `name`, `definition` (markdown, renders in the Definition tab)                                                                                      |
| Dashboard  | `name`, `url` (renders as a link in the Overview tab)                                                                                               |
| Chart      | `name`, `url` (renders as a link in the Overview tab)                                                                                               |


# Virtual Sources in Lineage

Virtual assets connect to any other asset in your lineage graph, filling gaps left by sources without native connectors.

Virtual assets participate in manual lineage exactly like native assets. You can draw lineage edges from a virtual asset to any other asset in your workspace — native or virtual — with no special handling required.

This lets you close gaps in your end-to-end lineage when a source doesn't have a native connector. For example, you can represent an on-premises database as a virtual source, create the relevant tables and columns, and link them into the lineage graph alongside your cloud warehouse assets.

<figure><img src="/files/n7q8EqbCPj093qqWXICI" alt=""><figcaption></figcaption></figure>

Adding lineage to a virtual asset follows the same process as any manual lineage addition. See [Add lineage relationships manually](/lineage/manual-lineage) for the full steps.


# Jira

Integrating Decube with Jira allows you to create issue tickets directly on the detected incidents, streamlining your workflow and ensuring efficient incident management.

## Prerequisites

1. An active Decube account with the necessary permissions to configure integrations.
2. A Jira account with administrative privileges or the required permissions to create and manage issues.

{% hint style="info" %}
The authentication credentials is different between Jira Cloud and Jira Data Centre. Please refer to the correct section for required step for each version.
{% endhint %}

## Set up Jira connection

1. Log in to your Decube account and navigate to the **Integrations** tab under **My** **Account** page.
2. Locate the **Third-party** tab on the right side and click on the **Connect to Jira** button, as shown below.

<figure><img src="/files/qlWAISnlbZuY20MGagXJ" alt=""><figcaption><p>Third-party integration</p></figcaption></figure>

3. Select your Infrastructure. There are 2 options, which depend on your organization's Jira setup: Cloud and Data centre.

### **Option 1: Jira Cloud**

<figure><img src="/files/2wiIVOVfEtTUNnuYxRoS" alt=""><figcaption><p>Cloud</p></figcaption></figure>

1. After selecting Jira Cloud option, you will need to provide the **Server URL**. This can be found in your workspace console settings, eg. `https://your-domain.atlassian.net`.
2. You will then need to generate an **API token** from your Atlassian account. This is the [Atlassian guide ](https://support.atlassian.com/atlassian-account/docs/manage-api-tokens-for-your-atlassian-account/)on how to create the API token.
3. Provide the **Jira account email** which was used to generate the API token above.

Once these information is provided, click on `Proceed to connect`

### **Option 2: Data Center**

<figure><img src="/files/kvAau4mbwHFeAzh85zdi" alt=""><figcaption><p>Data Center</p></figcaption></figure>

1. After selecting Jira Data Center option, you will need to provide the **Server URL**. This can be found in your workspace console settings, eg. `https://your-domain.atlassian.net`.
2. You will then need to generate an **Personal Access Token** from your Atlassian account. This is the [Atlassian guide](https://confluence.atlassian.com/enterprise/using-personal-access-tokens-1026032365.html) on how to create the API token.
3. Provide the **Jira Account Username** which was used to generate the Personal Access Token above.

Once these information is provided, click on `Proceed to connect`.

After successful authentication, you will see that the button to Connect to Jira now changes to: `Disconnect` and `Manage`.

<figure><img src="/files/qhR0fhbbpilm9ShgpN86" alt=""><figcaption><p>Jira Credentials Successfully Authenticated State.</p></figcaption></figure>

## Syncing Statuses from Decube to Jira

To Sync statuses from your Jira Project to Decube, Select the "Manage" option.

<figure><img src="/files/Gd4PC0qPco3LjKLAnqRA" alt=""><figcaption></figcaption></figure>

Add a project to begin mapping your Jira Statuses from the selected project to Decube's Statuses.

<figure><img src="/files/cgLjetdLzL9GsIZSVUl4" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/y3J64QL1l8KEwQRU9ynA" alt=""><figcaption></figcaption></figure>

Upon Selecting a project, select "+ Map Issue".

<figure><img src="/files/8Kkc9ULUEE7cgckK44wX" alt=""><figcaption></figcaption></figure>

Choose the Issue Type, Issue Status, and the corresponding Incident Status in Decube that you want to map.

<figure><img src="/files/P5b34bF0Tu4e60U33vPD" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/LwfjSqcSlm8MgV0K0wtm" alt=""><figcaption></figcaption></figure>

## Adding Jira Sync Webhook to Decube

To sync statuses from Jira to Decube and automatically close incidents, you will need to add following Jira Sync Webhook URL to your Jira account.

<figure><img src="/files/DIBhYGWX6Cnf0OHepryz" alt=""><figcaption></figcaption></figure>

The steps for managing webhooks in your Jira account differs from Jira Cloud and Jira Data Centre. Please refer to the correct guide for each Manage Webhook flow:

1. Jira Cloud: [link to Atlassian documentation](https://support.atlassian.com/jira-cloud-administration/docs/manage-webhooks/)
2. Jira Data Centre: [link to Atlassian documentation](https://developer.atlassian.com/server/jira/platform/webhooks/#registering-a-webhook)

{% hint style="info" %}
You will need to set "**Issues updated**" as the Jira event to be submitted to the Webhook endpoint.
{% endhint %}

## Creating Issues from Decube to Jira

To create an issue in Jira from Decube, go to the Data Quality module and select an incident. In the Incident Details Page, you will find an option to create a Jira issue.

<figure><img src="/files/X3T4FZ6BVp4Mni3ufVGw" alt=""><figcaption></figcaption></figure>

\
In the "Create a Jira Issue" pop-up, the user must fill out the following fields:

* Select a Project (based on the projects previously mapped [here](#syncing-statuses-from-decube-to-jira))
* Choose an Issue Type
* Add an Assignee
* Enter a Summary
* Provide a Description

<figure><img src="/files/NaFgTpTl2lWLqBtKm2KH" alt=""><figcaption></figcaption></figure>

Complete the process by selecting "Create This Issue." You have successfully created an issue from Decube to Jira. You can now view the following details:

* Issue Title
* Issue ID
* Assignee
* Issue Status
* Options to Unlink or Modify the Jira Issue.

<figure><img src="/files/l4KyjNjTBMFunaaRnf0v" alt=""><figcaption></figcaption></figure>

### Modifying a Jira Issue

To Modify a Jira issue, select the "Pencil" icon on the generated Jira Issue.

<figure><img src="/files/QE6q6KKdjnnUk54UeWZE" alt=""><figcaption></figcaption></figure>

Users are allowed to modify the fields below except Projects and Issue Type.

<figure><img src="/files/ZyWh9YgvCZxCGIlnqb30" alt=""><figcaption></figcaption></figure>

### Unlinking Jira Issues from Decube

To unlink a Jira Issue from Decube, select the "Unlink" option from the generated Jira Issue ticket.

<figure><img src="/files/xntykg6iu4yOVNv2VhxA" alt=""><figcaption></figcaption></figure>

On the confirmation prompt, select "Unlink" to unlink you Issue from Decube.

<figure><img src="/files/MLiAbQkmMCikXabYey0c" alt=""><figcaption></figcaption></figure>

When you unlink a Jira issue, Decube will stop syncing any update statuses for that issue. However, you can relink the issue at any time to resume syncing the latest status from Jira.

<figure><img src="/files/jrR3JseegLb9W5v3y82A" alt=""><figcaption></figcaption></figure>

### Disconnecting Jira from Decube

To Disconnect Jira, navigate to My Account > Integrations > Jira and simply select the "Disconnect" option.

<figure><img src="/files/sl1BN83OvKJc0ashXa9M" alt=""><figcaption></figcaption></figure>

Select "Yes, Disconnect" to Disconnect Jira from Decube.

<figure><img src="/files/O45gYogdeULx1xhXDY2q" alt=""><figcaption></figcaption></figure>

When you disconnect Jira from Decube, all generated Jira issue tickets will be revoked. However, you can reconnect to Jira at any time.

<figure><img src="/files/ece825vZ25arwAGfIrm5" alt=""><figcaption></figcaption></figure>

## FAQ

1. **I have marked "Done" on Jira, however it is not syncing to "Close" the incident on Decube.**

First, check that you have correctly mapped the Jira status correctly to the Decube "**Close**" status in the "**Manage Jira Connection**" as in section [#syncing-statuses-from-decube-to-jira](#syncing-statuses-from-decube-to-jira "mention"). You will need to make sure that the right Jira project/space is mapped correctly for which project/space you want to create the issue on.

If the above was done correctly, then you should check if the Webhook endpoint was set up correctly in Jira by following the section for [#adding-jira-sync-webhook-to-decube](#adding-jira-sync-webhook-to-decube "mention").


# Enabling VPC Access

In the situation that your organization's data sources is not publicly accessible, you will need to allow decube access to your data sources.

We support two access method depending on your cloud security configuration:

{% content-ref url="/pages/IeabSxBfHBW69C0lrCKD" %}
[IP Whitelisting](/security-and-connectivity/ip-whitelisting)
{% endcontent-ref %}

{% content-ref url="/pages/jU79kSLrPa7VgOtHKw6b" %}
[SSH Tunneling](/security-and-connectivity/ssh-tunneling)
{% endcontent-ref %}


# IP Whitelisting

If your access control is mainly based on ingress IP blocking, you may whitelist our IPs to allow our connectors to access your data sources.

{% hint style="info" %}
This is only applicable for customers on Multi-tenant deployments. If you are on Single-tenant setup, please contact your Account Manager for more information.
{% endhint %}

You only need to whitelist the IPs based on where your Decube SaaS instance is hosted.

**US-1:**

* 13.223.156.154/32
* 35.172.65.84/32
* 52.203.151.91/32

**APAC:**

* 52.77.53.119/32
* 52.76.27.251/32
* 18.141.95.11/32


# SSH Tunneling

You may opt to choose SSH Tunneling as well through a bastion host within your VPC to allow decube collectors access to your data source.

### SSH Bastion Host Setup

Setting up a SSH Bastion host differs based on the Cloud provider you are on. Here's some reference for the different providers:

* [AWS](https://aws.amazon.com/premiumsupport/knowledge-center/rds-connect-ec2-bastion-host/)
* [GCP](https://cloud.google.com/kubernetes-engine/docs/tutorials/private-cluster-bastion#create-bastion)
* [Azure](https://learn.microsoft.com/en-us/azure/bastion/bastion-overview)

### Getting your Public Key

When connecting to your data source for the first time, you can enable the "Enable SSH" toggle to get your organization-wide public key. Copy or download the file for later.

<figure><img src="/files/cBYYV2kkqmKsTlXfkaTl" alt=""><figcaption></figcaption></figure>

### Create A SSH User

We recommend creating a unique SSH User just for decube purposes. This lets you separate all configuration to only this unique user. To do so, access your own SSH host server as a privileged user or through sudo and run

```
sudo useradd -m [DECUBE_USERNAME] # choose any username (you'll need to pass this to us)
sudo passwd [DECUBE_USERNAME]     # you'll be prompted to input a password
                                  # this password is for your login only, do not share
                                  # with us
```

### Add Key to SSH Host

Access your own SSH host server as the previously created user and add the previously mentioned SSH Public Key to the `authorized_key` file. Usually this will be located at `/home/YOURUSER/.ssh/authorized_key`.

{% hint style="info" %}
You'd need to share with us the SSH username, SSH host and SSH port later
{% endhint %}


# AWS Identities

Create and use AWS Identities to connect across AWS data sources like AWS Glue and S3.

### Generating a Decube AWS Identity

* Step 1: Go to My Account → Integrations → Identities.
* Step 2: Click on **⊕ Add new**

{% hint style="info" %}
You can create up to 5 active Decube AWS Identities at one time.
{% endhint %}

<figure><img src="/files/qY9xTb8eXxQvj5jfBmRt" alt=""><figcaption></figcaption></figure>

* Step 3: Name the identity. This name will be shown in the Data Source form when adding/modifying connection to the data source.

<figure><img src="/files/tw40jomyKRksmlQkk8vv" alt=""><figcaption></figcaption></figure>

* Step 4: Take note of the **ARN** and **External ID**.

{% hint style="info" %}
The ARN and External ID are unique to your organisation. These values will be used when setting up the **AWS Role**.
{% endhint %}

<figure><img src="/files/TKwoZDxcd2mWrCHhS9v6" alt=""><figcaption><p>New AWS Identity was sucessfully added</p></figcaption></figure>

To understand how to use the AWS Identities in a specific data source connection, please go to the respective data sources page and see the section on **AWS Roles**. Link to the docs: [AWS Glue](/transformation-tools/aws-glue), [AWS S3](/datalake/s3) [dbt Core](/transformation-tools/dbt-core).

### Deleting a Decube AWS Identity

{% hint style="info" %}
Before you can delete a Decube AWS Identity, you need to ensure it is not being used as a connection option in any connected data sources.
{% endhint %}

<figure><img src="/files/B9fYX2eDX7ONzE1WVudZ" alt=""><figcaption><p>Example of an identity is being used by a data source</p></figcaption></figure>

* Step 1: Click on the bin icon on right side. Make sure the identity you want to delete is not being used by any connected data sources.

<figure><img src="/files/vdy3CzSJJYl2QPGXOEVg" alt=""><figcaption></figcaption></figure>

* Step 2: Enter "`Delete`" to confirm the deletion.

<figure><img src="/files/ba7OAilh4M2eaNNt1DtT" alt=""><figcaption></figcaption></figure>

* You will see a success message once the identity was successfully deleted.

<figure><img src="/files/kUwQq29Ti2AKjDcFcWKE" alt=""><figcaption></figcaption></figure>


# Incidents Overview

Monitor your data assets for quality issues and manage incidents across your organisation.

Decube's Data Quality module monitors your data assets continuously, detects anomalies automatically, and surfaces incidents so your team can investigate and resolve issues before they affect downstream consumers.

## Monitor Types

Each monitor type targets a specific category of data quality issue. Click through to the setup guide for the monitor you want to configure.

| Monitor          | What it detects                                              | Setup guide                                                                                              |
| ---------------- | ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------- |
| **Freshness**    | Data that has stopped updating within an expected window     | [Set up Freshness monitors](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)              |
| **Volume**       | Unexpected changes in row counts                             | [Set up Volume monitors](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)                    |
| **Field Health** | Column-level anomalies — nulls, uniqueness, ranges, patterns | [Set up Field Health monitors](/data-quality/how-to-set-up-monitors/set-up-field-tests)                  |
| **Schema Drift** | Table structure changes                                      | [Set up Schema Drift monitors](/data-quality/how-to-set-up-monitors/set-up-schema-drift-monitors)        |
| **Custom SQL**   | Business rule violations defined by custom SQL logic         | [Set up Custom SQL monitors](/data-quality/how-to-set-up-monitors/custom-sql-monitors)                   |
| **Job Failure**  | ETL pipeline job execution failures                          | [Set up Job Failure monitors](/data-quality/how-to-set-up-monitors/set-up-data-job-job-failure-monitors) |
| **Grouped-By**   | Quality issues segmented by dimension values                 | [Set up Grouped-By monitors](/data-quality/how-to-set-up-monitors/set-up-grouped-by-monitors)            |

Monitors run in two modes: **Scheduled** (continuous, at a configurable frequency) and **On-Demand** (manual, for ad-hoc validation).

## Getting Started

1. [Enable asset monitoring](/data-quality/enable-asset-monitoring) on the tables you want to cover.
2. [Set up alert notifications](/alert-notifications/notification-alerts) so your team is notified when incidents are triggered.
3. Review [Incident Details](/data-quality/data-quality/incident-details) to understand what each incident shows you and how to investigate.
4. Review [Managing Incidents](/data-quality/data-quality/incident-management) to learn how to close, mute, and bulk-update incidents.

{% content-ref url="/pages/f7SehvK9kQi9RYPSSxru" %}
[Available Monitor Types](/data-quality/available-monitor-types)
{% endcontent-ref %}

## Related Pages

{% content-ref url="/pages/JwDGUfWup3P1TQWUn9zw" %}
[Incident Details](/data-quality/data-quality/incident-details)
{% endcontent-ref %}

{% content-ref url="/pages/ZayzgoP6LWwmhawZ0z1U" %}
[Managing Incidents](/data-quality/data-quality/incident-management)
{% endcontent-ref %}

{% content-ref url="/pages/cFfu9PDAnjRwAkZxU098" %}
[Config Settings](/data-quality/config-settings)
{% endcontent-ref %}

{% content-ref url="/pages/oHQMBTOmZQevImaIjmyw" %}
[Incident model feedback](/data-quality/data-quality/incident-model-feedback)
{% endcontent-ref %}

{% content-ref url="/pages/AuBPzQbjZY9jvr7I4oRR" %}
[Asset Report: Data Quality Scorecard](/reports/asset-report-data-quality-scorecard)
{% endcontent-ref %}


# Incident Details

Explore the Incident Details page to investigate root causes, review history, and assess downstream impact.

When you select any incident from the Data Quality module, you are taken to the **Incident Details** page. This page gives you the full context of the incident: what triggered it, the historical trend, who owns it, and which downstream assets are at risk.

## Incident Types

Decube categorises incidents into six types based on the monitor that triggered them.

| Type             | What triggered it                                | Common use case                                     |
| ---------------- | ------------------------------------------------ | --------------------------------------------------- |
| **Freshness**    | Data stopped updating within the expected window | Real-time dashboards and daily reports              |
| **Volume**       | Row count changed unexpectedly                   | Missing data loads or pipeline issues               |
| **Field Health** | Column-level anomaly detected                    | Null spikes, uniqueness violations, range breaches  |
| **Schema Drift** | Table structure changed                          | Preventing downstream application failures          |
| **Custom SQL**   | A custom SQL validation rule was breached        | Complex business logic and data relationship checks |
| **Job Failure**  | An ETL pipeline job failed                       | Data transformation process monitoring              |

## Assignee & Audit History

You can assign any incident to a team member directly from the Incident Details page. Open the incident, click the **Assignee** field in the right panel, and select the relevant person.

All actions taken on an incident — assignments, status changes, comments — are logged automatically in the **Audit History** section at the bottom right of the page.

<figure><img src="/files/ypkaDyjbByXiF22OKbvB" alt="The incident details page showing the Add Assignee panel and Audit History section"><figcaption><p>Assignee management and Audit History on the Incident Details page</p></figcaption></figure>

## Tabs

The Incident Details page is organised into tabs. Each tab surfaces a different dimension of the incident.

### Details

The Details tab shows the core metadata for the incident: the affected asset, incident type, level, monitor mode, test name, and description. This is the starting point for understanding what triggered the alert and what value was observed.

### Preview

The Preview tab shows a data preview of the affected table or column at the time the incident was detected.

### History

The History tab lists past monitor runs for the affected asset, including the metrics from each scan. Use it to identify the run that first surfaced the anomaly and compare it against successful scans to understand what changed.

<figure><img src="/files/iWvRXRwoFh6f9C9ux0Vp" alt="The History tab showing a list of past monitor runs with pass and fail statuses"><figcaption><p>The History tab for a Field Health incident</p></figcaption></figure>

### Impacted Areas

The Impacted Areas tab generates a list of downstream assets that may be affected by the incident, based on lineage. This includes downstream tables, jobs, and dashboards. You can export the list as a CSV and share it with the owners of those assets.

<figure><img src="/files/zaLSuOtcthr6cxoaXZ7S" alt="The Impacted Areas tab listing downstream tables and dashboards affected by the incident"><figcaption><p>Downstream assets impacted by a Custom SQL incident</p></figcaption></figure>

<figure><img src="/files/DSVB1YvQ4Ov40JYscKfN" alt="Continued list of impacted downstream assets"><figcaption><p>Continued list of impacted assets</p></figcaption></figure>

### Lineage

The Lineage tab surfaces column-level lineage directly within the Incident Details page, so you can trace the root cause and downstream impact of an incident without leaving the incident context.

<figure><img src="/files/UlCHsY8G2DCPUq4RR9xX" alt=""><figcaption></figcaption></figure>

When you open the Lineage tab, Decube renders a column-level lineage graph centred on the affected asset. The graph shows the upstream sources feeding into the affected column and the downstream tables, jobs, and dashboards that depend on it. You can zoom, pan, and reset the view using the controls in the graph toolbar.

* **Trace the source of bad data** — follow upstream connections to identify where an anomalous value or null originated.
* **Assess downstream exposure** — see which tables and BI assets consume the affected column, so you can prioritise fixes by blast radius.
* **Inspect related columns** — each node in the graph shows the related columns for that asset, giving you the full propagation path at a glance.

{% hint style="info" %}
The lineage graph reflects the connections Decube has detected for that asset.
{% endhint %}

{% content-ref url="/pages/ZRdnhEirjf6oOvayCLu8" %}
[Automated Lineage](/lineage/automated-lineage)
{% endcontent-ref %}

## Related Pages

{% content-ref url="/pages/ZayzgoP6LWwmhawZ0z1U" %}
[Managing Incidents](/data-quality/data-quality/incident-management)
{% endcontent-ref %}

{% content-ref url="/pages/oHQMBTOmZQevImaIjmyw" %}
[Incident model feedback](/data-quality/data-quality/incident-model-feedback)
{% endcontent-ref %}


# Managing Incidents

Close, mute, and bulk-update incidents to manage data quality issues across your organisation.

When a monitor detects an anomaly, Decube opens an incident with a status of `open`. From there, you can close it once resolved, mute it to suppress repeated alerts, or update multiple incidents at once using bulk actions.

## Incident Status

Every incident has one of three statuses: **Open**, **Closed**, or **Muted**.

* **Close** an incident when the underlying issue has been resolved.
* **Mute** an incident for a set time period to prevent duplicate alerts while you investigate. Muting is useful when another incident on the same table or column is already being tracked. Decube automatically unmutes the incident after the period you select.

To update an incident's status, open the incident and use the status controls in the right panel.

<figure><img src="/files/3ESNOeV4LEIbS1KGZNBy" alt="The right panel of an open incident showing the status control options"><figcaption><p>Status controls on an open incident</p></figcaption></figure>

{% hint style="info" %}
To filter the Incidents Overview by status, click **Apply Filters** and select the relevant checkboxes under **Incident Status**.
{% endhint %}

<figure><img src="/files/WftTo27ArL6oV7T44zur" alt="The Apply Filters modal with Incident Status checkboxes"><figcaption><p>Filtering incidents by status</p></figcaption></figure>

## Bulk Update Incident Status

You can update multiple incidents at once directly from the Incidents Overview page, without opening each one individually.

**Selecting incidents**

A checkbox column appears at the left of the incidents table. Check individual rows to select specific incidents, or use the header checkbox to select all incidents currently loaded on the page. Selected rows are highlighted, and the top-left of the table shows a running count — for example, `3 incidents selected`. Use **Clear selection** to deselect everything and start again.

Filtering or searching within the page preserves your current selection unless you navigate away from the Incidents Overview page.

{% hint style="info" %}
Incidents from assets you don't have edit access to display a lock icon instead of a checkbox and cannot be selected for bulk actions.
{% endhint %}

**Performing a bulk action**

Once at least one editable incident is selected, the **Perform bulk action** button activates. Click it to open the bulk update modal and choose a target status:

* Unmute incident
* Close incident
* Mute for 1 day
* Mute for 1 week
* Mute for 1 month

Before confirming, Decube shows a preview of exactly what will happen to each selected incident, including which will be updated and which will be skipped:

* Incidents already in the target state are skipped automatically. The exception is "Muted to Muted" transitions — these reset the mute duration to the newly selected period.
* Closed incidents cannot be re-opened or muted via bulk action.
* The **Confirm changes** button only activates when at least one incident will actually change state.

<figure><img src="/files/CVdidmF6pgAGI1KiBggG" alt="The bulk update modal showing a preview of status changes before confirmation"><figcaption><p>Bulk update preview modal</p></figcaption></figure>

{% hint style="warning" %}
Bulk status updates cannot be undone once confirmed.
{% endhint %}

**Traceability**

Each incident updated via a bulk action has the change recorded in its audit history with a `via bulk action` label.

**Limits**

* A maximum of 1,000 incidents can be selected at one time.
* The selection resets when you navigate away from the Incidents Overview page.

## Related Pages

{% content-ref url="/pages/JwDGUfWup3P1TQWUn9zw" %}
[Incident Details](/data-quality/data-quality/incident-details)
{% endcontent-ref %}

{% content-ref url="/pages/oHQMBTOmZQevImaIjmyw" %}
[Incident model feedback](/data-quality/data-quality/incident-model-feedback)
{% endcontent-ref %}


# Incident model feedback

How to adjust a monitor's alert sensitivity using the feedback mechanism on an incident.

For monitors that use Smart Training, you can adjust alert sensitivity directly from an incident. Sensitivity controls how wide or narrow the model's confidence interval is — a wider interval means fewer alerts; a narrower interval means more.

Sensitivity is set on a scale of **–5 to +5**:

| Slider position | Confidence interval | Effect                                  |
| --------------- | ------------------- | --------------------------------------- |
| –5              | 0.99                | Widest — least sensitive, fewest alerts |
| 0 (default)     | 0.90                | Balanced                                |
| +5              | 0.80                | Narrowest — most sensitive, most alerts |

Each step on the slider changes the confidence interval by 0.02.

***

## How to provide feedback

1. Navigate to **Data Quality > Incidents** and open the incident you want to provide feedback on.
2. Click the thumbs up (👍) icon if the alert was accurate. Click the thumbs down (👎) icon if it was a false positive or missed something.
3. For monitors with Smart Training enabled, clicking thumbs down reveals the **sensitivity slider**. Adjust the slider and submit — the new sensitivity takes effect on the next scan.

{% hint style="info" %}
Sensitivity adjustment is available for Freshness, Volume, and Field Health monitors with Smart Training (Auto threshold) enabled. It is not available for Custom SQL monitors or monitors using manual thresholds.
{% endhint %}

***

## When to adjust

* **Too many false positives** → move the slider left (toward –5) to widen the confidence interval.
* **Missing anomalies you expect to catch** → move the slider right (toward +5) to narrow the interval.

For newly created monitors, wait until at least a few weeks of scan history have accumulated before adjusting sensitivity — early scans may not yet reflect the true variance in your data.


# How Anomaly Detection Works

How Decube detects anomalies in your data — what a monitor is, how the ML model trains, and what determines whether an alert fires.

Understanding how Decube's anomaly detection operates helps you configure monitors that behave predictably and interpret incidents accurately.

## What a monitor is

A monitor represents one test applied to one asset. Each monitor produces its own independent incident stream — if you apply three different tests to the same table column, you have three monitors, each of which can open and close incidents independently.

## How the ML model learns

For monitors that use **Smart Training**, Decube runs an ML model against the asset's historical data to learn the normal range for a given metric. The model builds a confidence interval — an expected upper and lower bound — for each scan point. When a new scan falls outside that interval, Decube opens an incident.

The confidence interval widens or narrows based on the **Sensitivity** setting you choose. See [Sensitivity](#sensitivity) below.

### Historical lookback by scan frequency

When a new monitor is created (or retrained), the model collects historical data to train on. The amount of history collected depends on the scan frequency you configure:

| Scan frequency | Historical lookback |
| -------------- | ------------------- |
| Hourly         | 7 days              |
| Every 6 hours  | 30 days             |
| Every 12 hours | 60 days             |
| Daily          | 192 days            |
| Weekly         | 395 days            |

Monitors scan on a schedule after training is complete. During the training period, the monitor is visible in **All Monitors** but does not produce incidents.

### Sparse data and silent skipping

The ML model requires a minimum amount of valid signal before it can produce a reliable confidence interval. If a scan finds fewer than **5 valid data points in the last 30 observations**, the model will not run the test based on the collected metrics yet until the threshold is set.

This is intentional behaviour: firing an alert on insufficient data would produce unreliable signals. However, it means a monitor on a low-volume or infrequently-updated table may appear inactive. If your monitors are not producing incidents on a table you expect to have anomalies, check whether the table has enough scan history to meet the threshold.

{% hint style="warning" %}
Monitors on sparse tables can train and appear healthy while silently skipping every scan. If you need coverage on a low-volume table, consider On-Demand mode with a manual threshold instead of Smart Training.
{% endhint %}

## Sensitivity

The Sensitivity setting controls how wide or narrow the model's confidence interval is, which in turn controls how easy it is for a data point to fall outside it and trigger an incident.

The scale runs from **–5 to +5**:

| Value | Effect                                                         |
| ----- | -------------------------------------------------------------- |
| –5    | Widest confidence interval — least sensitive, fewest incidents |
| 0     | Default — balanced sensitivity                                 |
| +5    | Narrowest confidence interval — most sensitive, most incidents |

You set Sensitivity via the feedback mechanism on an incident.

{% content-ref url="<https://github.com/DecubeIO/decube-docs/blob/public/incident-model-feedback.md>" %}
<https://github.com/DecubeIO/decube-docs/blob/public/incident-model-feedback.md>
{% endcontent-ref %}

## Monitors that do not use the ML model

Not all monitor types use the ML model. Monitors with a fixed threshold configuration (Absolute, Percentage, Positive Range, or Any Range) compare each scan result directly against the bounds you define — no training period, no confidence interval.

For these monitors, the sparse-data rule and training timeout do not apply.

{% content-ref url="/pages/JO3VkDrrA78BnGdP7coz" %}
[Setting Up Your Data Quality Thresholds](/data-quality/monitor-configuration-settings/setting-up-your-data-quality-thresholds)
{% endcontent-ref %}


# Enable asset monitoring

How to create, view, and manage monitors for your data assets from the Config module.

The **Config** module is where you create and manage all data quality monitors. It has two tabs: **All Monitors**, which shows every monitor across your connected sources, and **Create**, where you set up new monitors.

{% embed url="<https://youtu.be/fJ7SvrkE6Mo?si=nnvUXWNnfp3RrNNu>" %}

Before creating your first monitor, make sure your data source is connected and your alert channels are configured in [Config Settings](/data-quality/config-settings).

{% hint style="info" %}
Grouped-By monitors are set up within each monitor type's creation flow. Schema Drift and Job Failure monitors are managed directly from the All Monitors tab.
{% endhint %}

***

## All Monitors tab <a href="#all-monitors-tab" id="all-monitors-tab"></a>

The **All Monitors** tab lists every monitor in your organisation, with its current status, last run time, and scan frequency.

<figure><img src="/files/4tNEIc92optPOsnXIvvq" alt=""><figcaption><p>All Monitors tab</p></figcaption></figure>

To manage an existing monitor, click the ellipsis (︙) at the far right of the monitor row. This opens two options:

**View Monitor** — opens the monitor configuration form, where you can modify settings, disable the monitor, or delete it.

<figure><img src="/files/VRgDL3d51VKLHCti3Tlt" alt=""><figcaption><p>Monitor configuration form</p></figcaption></figure>

<figure><img src="/files/GHiQBKRNFBitkD7CBokb" alt=""><figcaption><p>Delete monitor option</p></figcaption></figure>

**Monitor Info** — opens the monitor's history and performance data, with two tabs:

* **History** — lists each scan run with its timestamp and result status:
  * **Passed** — the scan ran and the metric was within the expected range.
  * **Failed** — the scan ran and the metric fell outside the expected range, triggering an incident.
  * **Skipped** — the scan ran but could not collect metrics, or the baseline threshold has not yet been established. See [sparse data behaviour](/data-quality/anomaly-detection-explained#sparse-data-and-silent-skipping).
  * **Errored** — the scan could not run. Contact <support@decube.io> if this persists.

<figure><img src="/files/LKgYRzCasD4lt5C40k7S" alt=""><figcaption><p>Monitor scan history</p></figcaption></figure>

* **Performance** — shows the monitored metric over time, alongside the confidence interval or threshold bounds.

<figure><img src="/files/ZB9HiC59P0MqllUsFtwE" alt=""><figcaption><p>Monitor performance view</p></figcaption></figure>

***

## Create tab

Navigate to **Config > Create** to set up a new monitor. Select the monitor type card that matches what you want to detect.

<figure><img src="/files/df7iQGWIvY91Y1DM6fOC" alt=""><figcaption><p>Monitor type cards in the Create tab</p></figcaption></figure>

### Available monitor types

| Monitor type                                                                             | What it detects                                              | Setup guide                                                                                              |
| ---------------------------------------------------------------------------------------- | ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------- |
| [Freshness](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)              | Data that has stopped arriving on its expected schedule      | [Set up Freshness monitors](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)              |
| [Volume](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)                    | Row-count increments outside the expected range              | [Set up Volume monitors](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)                    |
| [Field Health](/data-quality/how-to-set-up-monitors/set-up-field-tests)                  | Column-level anomalies — nulls, uniqueness, ranges, patterns | [Set up Field Health monitors](/data-quality/how-to-set-up-monitors/set-up-field-tests)                  |
| [Custom SQL](/data-quality/how-to-set-up-monitors/custom-sql-monitors)                   | Business rule violations defined by custom SQL logic         | [Set up Custom SQL monitors](/data-quality/how-to-set-up-monitors/custom-sql-monitors)                   |
| [Schema Drift](/data-quality/how-to-set-up-monitors/set-up-schema-drift-monitors)        | Table structure changes (auto-enabled)                       | [Manage Schema Drift monitors](/data-quality/how-to-set-up-monitors/set-up-schema-drift-monitors)        |
| [Job Failure](/data-quality/how-to-set-up-monitors/set-up-data-job-job-failure-monitors) | ETL pipeline job execution failures (auto-enabled)           | [Manage Job Failure monitors](/data-quality/how-to-set-up-monitors/set-up-data-job-job-failure-monitors) |

For a full description of each type, see [Available Monitor Types](/data-quality/available-monitor-types).

***

## Getting started

If you're setting up monitoring for the first time:

1. Start with [Freshness](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors) and [Volume](/data-quality/how-to-set-up-monitors/set-up-volume-monitors) monitors on your most business-critical tables.
2. Add [Field Health](/data-quality/how-to-set-up-monitors/set-up-field-tests) monitors on key columns once the table-level monitors are in place.
3. Use [Custom SQL](/data-quality/how-to-set-up-monitors/custom-sql-monitors) for any validation logic that the built-in tests don't cover.

Schema Drift and Job Failure monitors are enabled automatically when you connect a data source — no setup required.


# Config Settings

How to configure default alert channels and default monitoring behaviour for your organisation.

The **Config Settings** page has two tabs: **Config Alerts** for setting up notification channels, and **Default Settings** for defining the default scan frequency and incident severity applied to new monitors.

{% embed url="<https://www.loom.com/share/60938e882ed94612b8c41a82acbb8f7d?sid=3eea07df-33ae-4979-a1c1-1dc9a2cf8da6>" %}

To access Config Settings, go to **My Account** and select the **Config Settings** tab.

***

## Config Alerts

The **Config Alerts** tab is where you set up your organisation's default notification channels — the email addresses and Slack channels that receive alerts when a monitor opens an incident.

<figure><img src="/files/MLiCCTa7SIJqw7bBezEF" alt=""><figcaption><p>Config Alerts tab</p></figcaption></figure>

For step-by-step instructions on connecting each alert channel:

{% content-ref url="/pages/QEgoy1vIsDVzEgiD9SEb" %}
[Get alerts on email](/alert-notifications/notification-alerts)
{% endcontent-ref %}

***

## Default Settings

The **Default Settings** tab lets you define organisation-wide defaults for two monitor properties:

| Setting                  | What it controls                                                                                                                                           |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Monitoring Frequency** | The default scan frequency applied when a new monitor is created (for example, Daily). Individual monitors can override this at creation time.             |
| **Incident Severity**    | The default incident level (Critical, High, Medium, or Low) assigned to incidents on new monitors. Individual monitors can override this at creation time. |

<figure><img src="/files/gZeCMnJwrkZGY1oNLD7v" alt=""><figcaption><p>Default Settings tab</p></figcaption></figure>

{% hint style="warning" %}
Default Settings apply only to **newly added data sources and monitors**. Changing defaults here does not affect the frequency or severity of existing monitors. To update an existing monitor's frequency or severity, edit the monitor directly from **All Monitors**.
{% endhint %}

{% hint style="info" %}
Default Settings do not apply to Grouped-By monitors. Grouped-By monitors must have frequency and severity configured individually at creation time.
{% endhint %}


# Available Monitor Types

The monitor types available in Decube's Data Quality module and what each one detects.

{% embed url="<https://youtu.be/k9fM4a8KICE>" %}

Decube provides six monitor types, each designed to detect a different category of data quality issue. Schema Drift and Job Failure monitors are enabled automatically when you connect a data source. All other types are created manually from the Config module.

***

## Table-level monitors

### Schema Drift

Schema Drift monitors detect structural changes to your tables — column additions, deletions, data type changes, and other schema modifications. They are enabled automatically on all tables when you connect a data source and require no additional setup.

{% content-ref url="/pages/BK77u7wZMk210T4bia4f" %}
[Modify Schema Drift Monitors](/data-quality/how-to-set-up-monitors/set-up-schema-drift-monitors)
{% endcontent-ref %}

### Freshness

Freshness monitors detect when data stops arriving on its expected schedule. The scheduled Freshness monitor uses an ML model trained on historical arrival patterns — it learns when data normally shows up and raises an incident when the probability of arrival drops below 50%. This means it accounts for expected quiet periods (weekends, off-hours) without alerting unnecessarily.

{% content-ref url="/pages/YpzBR4Fms5X6aZfHSJXj" %}
[Set Up Freshness Monitors](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)
{% endcontent-ref %}

### Volume

Volume monitors detect when the row-count increment added to a table per scan window falls outside the expected range. The monitor tracks the number of new rows added since the last scan — not the total row count — so it catches both missing data loads and unexpected data spikes.

{% content-ref url="/pages/K9nwiEHp0zGa6BN2j0kl" %}
[Set Up Volume Monitors](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)
{% endcontent-ref %}

***

## Column-level monitors

### Field Health

Field Health monitors run tests against individual columns — null rates, uniqueness, value ranges, string patterns, and more. Available test types include:

* **Null** — monitors null values (Absolute, Percentage, or Auto threshold)
* **Unique** — monitors duplicate values (Absolute, Percentage, or Auto threshold)
* **Average / Min / Max** — monitors statistical metrics (Range or Auto threshold)
* **Cardinality** — tracks distinct value counts using a 5-period rolling window (Scheduled only; no training period required)
* **String Length** — validates string lengths (Range or Auto threshold)
* **Email / UUID / Regex Match** — validates format compliance (Absolute, Percentage, or Auto threshold)

{% hint style="info" %}
Available test types depend on the column's data type. For example, Average, Min, and Max apply to numeric columns only.
{% endhint %}

{% content-ref url="/pages/Dtvx04oJ0XpeTvHOkQQD" %}
[Set Up Field Health Monitors](/data-quality/how-to-set-up-monitors/set-up-field-tests)
{% endcontent-ref %}

***

## Advanced monitors

### Custom SQL

Custom SQL monitors let you define validation logic in SQL. You write a query that returns the failing rows; Decube counts those rows and opens an incident when the count is greater than zero. Custom SQL monitors attach to a data source connection rather than a specific table, so your query can reference any table in that source.

Custom SQL monitors do not support Smart Training. All thresholds are set manually.

{% content-ref url="/pages/PWGhO7SkX4GtJu3jfRBX" %}
[Set Up Custom SQL Monitors](/data-quality/how-to-set-up-monitors/custom-sql-monitors)
{% endcontent-ref %}

### Job Failure

Job Failure monitors detect failed runs in connected ETL tools (dbt, Airflow, and others). They are auto-created when you connect an ETL-type data source and require no additional configuration.

{% content-ref url="/pages/QRC6dLfH1ui8linyrJWK" %}
[Modify Job Failure Monitors (Data Job)](/data-quality/how-to-set-up-monitors/set-up-data-job-job-failure-monitors)
{% endcontent-ref %}

***

## Grouped-By monitors

Any monitor type (except Volume and Freshness) can be configured with a Group By column, creating one sub-monitor per distinct value in that column. Each sub-monitor tracks its own metric and opens incidents independently, giving you per-segment visibility without creating monitors manually for each value.

{% content-ref url="/pages/6W2WRDo4OLRfQIOt0QLk" %}
[Grouped-by Monitors](/data-quality/how-to-set-up-monitors/set-up-grouped-by-monitors)
{% endcontent-ref %}


# Monitor Configuration Settings

Reference pages for every monitor configuration setting.

1. [Monitor Configuration Reference](/data-quality/monitor-configuration-settings/configuration-reference) — every setting, what it does, and which are locked after creation
2. [Available Monitor Modes](/data-quality/monitor-configuration-settings/available-monitor-modes) — Scheduled vs On-Demand
3. [Custom Scheduling For Monitors](/data-quality/monitor-configuration-settings/custom-scheduling-for-monitors) — scan frequency options
4. [Setting Up Your Data Quality Thresholds](/data-quality/monitor-configuration-settings/setting-up-your-data-quality-thresholds) — threshold types and DQ score
5. [Retraining Monitors](/data-quality/monitor-configuration-settings/retraining-monitors) — what triggers a retrain and what data gets deleted


# Monitor Configuration Reference

A reference for every setting available when creating or editing a monitor, including which settings are locked after creation.

This page describes every setting available in the monitor creation and edit forms. Use it to understand what each setting does, what values are valid, and which settings you cannot change after a monitor is created.

## Settings locked after creation

The following settings cannot be changed on an existing monitor. To change them, delete the monitor and create a new one.

| Setting                                  | Why it's locked                                                                                        |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------------ |
| **Asset** (Source, Schema, Dataset)      | The monitor is bound to a specific table. Changing the asset would invalidate all historical data.     |
| **Test type**                            | Each test type measures a different metric. Changing it would make historical comparisons meaningless. |
| **Monitor mode** (Scheduled / On-Demand) | The training approach and data collection method differ fundamentally between modes.                   |
| **Group by**                             | Group-by configuration determines how the monitor partitions its data at creation time.                |

***

## All settings

### Asset

The table the monitor runs against. You select a **Source**, optionally filter by **Schema**, then select a **Dataset** (table). This setting is locked after creation.

***

### Test type

The metric the monitor measures (for example, Null, Freshness, Volume, Cardinality). Each test type has a fixed set of compatible threshold types and modes. Test type is locked after creation.

See [Set Up Field Health Monitors](/data-quality/how-to-set-up-monitors/set-up-field-tests) for the full list of Field Health test types and their compatibility.

***

### Monitor mode

Whether the monitor runs on a schedule (**Scheduled**) or only when manually triggered (**On-Demand**). Monitor mode is locked after creation.

| Mode          | When to use                                                |
| ------------- | ---------------------------------------------------------- |
| **Scheduled** | Production monitoring, continuous coverage, Smart Training |
| **On-Demand** | Ad-hoc validation, development and testing, manual runs    |

{% hint style="info" %}
On-Demand mode does not support Smart Training or Group By. See [Available Monitor Modes](/data-quality/monitor-configuration-settings/available-monitor-modes) for the full comparison.
{% endhint %}

***

### Row filtering

Controls which rows Decube considers when calculating the monitored metric. Three options are available:

| Option             | Behaviour                                                                                                                                                                                                                                                 |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Timestamp**      | Decube uses the selected timestamp column to identify rows added since the last scan. This is the most accurate option for incremental tables.                                                                                                            |
| **SQL Expression** | You provide an expression (in your data source's SQL dialect) that produces a timestamp. Use this when timestamps are stored in non-standard formats (strings, Unix epoch, split date/time columns). Validating the expression before saving is required. |
| **All Records**    | Decube scans the entire table on every run. No timestamp column is needed, but this mode cannot track increments reliably on tables where rows are deleted.                                                                                               |

{% hint style="warning" %}
**All Records does not support Smart Training.** If you select All Records as the row filtering mode, the Smart Training toggle is unavailable. The monitor will use a manual threshold instead of the ML model.
{% endhint %}

Changing the row filtering mode (or the timestamp column / SQL expression) after creation triggers a retrain and deletes all historical data. See [Retraining Monitors](/data-quality/monitor-configuration-settings/retraining-monitors).

***

### Smart Training

When enabled, Decube trains an ML model on historical data to establish a dynamic confidence interval. Smart Training requires:

* A timestamp-based row filtering mode (Timestamp or SQL Expression)
* Scheduled monitor mode (not On-Demand)

When Smart Training is enabled, the **Lookback Period** becomes configurable and a training period begins. See [How Anomaly Detection Works](/data-quality/anomaly-detection-explained) for details on the training process, historical lookback windows, and sparse-data behaviour.

Toggling Smart Training on after creation triggers a retrain.

***

### Scan frequency

How often a Scheduled monitor runs. Options range from hourly to monthly. The frequency you choose affects how much historical data the ML model collects during training.

See [Custom Scheduling for Monitors](/data-quality/monitor-configuration-settings/custom-scheduling-for-monitors) for all available frequency options.

{% hint style="warning" %}
Changing scan frequency on a monitor that uses Smart Training triggers a retrain and deletes all historical data. See [Retraining Monitors](/data-quality/monitor-configuration-settings/retraining-monitors).
{% endhint %}

***

### Threshold type

The method used to evaluate whether a scan result is anomalous. Four types are available:

| Threshold type     | How it works                                                                        | Compatible test types                                                    |
| ------------------ | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| **Absolute**       | Compares the raw row count of failing records against a min/max bound               | Null, Unique, Email, UUID, Regex Match                                   |
| **Percentage**     | Compares the percentage of failing records against a min/max bound (0–100)          | Null, Unique, Email, UUID, Regex Match                                   |
| **Positive Range** | Compares a numeric metric against a min/max bound (values ≥ 0)                      | Average, Min, Max, String Length                                         |
| **Any Range**      | Compares a numeric metric against a min/max bound (negative values allowed)         | Average, Min, Max                                                        |
| **Auto**           | ML model sets the bounds dynamically based on historical data. Scheduled mode only. | Null, Unique, Email, UUID, Regex Match, Average, Min, Max, String Length |

At least one bound (min or max) is required for manual threshold types. Leave the other bound empty to create a one-sided check (e.g., max only = "alert if the value exceeds X").

Threshold type is safe to change after creation — it does not trigger a retrain.

***

### Threshold bounds

The min and/or max values that define the acceptable range for the monitored metric. Bounds are always interpreted relative to the selected threshold type:

* For **Absolute**: non-negative integers (row counts)
* For **Percentage**: integers between 0 and 100
* For **Positive Range**: non-negative numeric values
* For **Any Range**: any numeric value including negatives

Threshold bounds are safe to change after creation — they do not trigger a retrain.

***

### Group by

Splits the monitor into one sub-monitor per distinct value in the selected column. Each sub-monitor tracks its own metric and opens incidents independently.

{% hint style="warning" %}
Group by is limited to **100 distinct values** per column. Columns with more than 100 distinct values are not eligible. See [Grouped-By Monitors](/data-quality/how-to-set-up-monitors/set-up-grouped-by-monitors).
{% endhint %}

Group by is locked after creation and is not available in On-Demand mode.

***

### Lookback period

Available on On-Demand monitors only. Defines how far back in time the monitor scans when it runs. For example, a lookback period of 7 days checks for qualifying rows from the past 7 days.

On Scheduled monitors with Smart Training enabled, the lookback period is set automatically based on scan frequency — see the [historical lookback table](/data-quality/anomaly-detection-explained#historical-lookback-by-scan-frequency).

***

### Incident level

The severity assigned to incidents this monitor opens. Options are typically **Critical**, **High**, **Medium**, and **Low**. Incident level affects notification routing if you have level-based alert rules configured.

Incident level is safe to change after creation — it does not trigger a retrain.

***

### Notifications

Controls which channels receive alerts when this monitor opens an incident. Toggle **Notify default channel** to enable custom routing, then add email addresses or Slack channel names.

If left off, incidents are still created and visible in the Data Quality module but no external notification is sent.

Notification settings are safe to change after creation — they do not trigger a retrain.

***

### Quality Dimension

An optional classification that maps the monitor to a data quality dimension (for example, Completeness, Validity, Timeliness). Dimension affects how incidents are grouped in DQ score reporting.

Quality Dimension is safe to change after creation.

***

### Name and Description

A human-readable name and optional description for the monitor. You can create multiple monitors of the same test type on the same column — use names to distinguish them.

Both fields are safe to change after creation — they do not trigger a retrain.


# Available Monitor Modes

Scheduled mode runs monitors automatically at a fixed frequency. On-Demand mode runs monitors manually when you choose.

Every monitor runs in one of two modes: **Scheduled** or **On-Demand**. You select the mode when creating a monitor, and it cannot be changed after creation.

## Scheduled

Scheduled monitors run automatically at the frequency you configure — hourly, daily, weekly, or custom intervals. Decube scans the asset at each interval, evaluates the monitored metric against the threshold or confidence interval, and opens an incident if the result falls outside the expected range.

Scheduled mode supports Smart Training. When Smart Training is enabled, the ML model trains on historical data to learn a dynamic confidence interval. See [How Anomaly Detection Works](/data-quality/anomaly-detection-explained) for how training works.

**All monitor types support Scheduled mode** — Freshness, Volume, Field Health (including Cardinality), Custom SQL, Schema Drift, Job Failure.

***

## On-Demand

On-Demand monitors do not run on a schedule. You trigger them manually from **All Monitors**, either immediately after creation (Save and Run) or at any later point (Run once from the ellipsis menu).

On-Demand mode is useful for ad-hoc investigation, validating a specific data load, or testing a monitor's configuration before committing to a schedule.

**Constraints in On-Demand mode:**

* Smart Training is not available. Thresholds must be set manually.
* Grouped By is not available.
* Frequency is not configured — instead, you set a **Lookback Period** that defines how far back each run scans.
* The Cardinality test type is not available.

**Supported monitor types:** Freshness, Volume, Field Health (excluding Cardinality), Custom SQL.

***

## Comparison

|                    | Scheduled                         | On-Demand                |
| ------------------ | --------------------------------- | ------------------------ |
| Runs automatically | Yes                               | No — manual trigger only |
| Smart Training     | Yes (with Timestamp row creation) | No                       |
| Grouped By         | Yes                               | No                       |
| Cardinality        | Yes                               | No                       |
| Frequency setting  | Required                          | Not applicable           |
| Lookback Period    | Set automatically from frequency  | Required                 |

***

## Choosing a mode

Use **Scheduled** for production monitoring — tables that need continuous coverage, SLA tracking, or anomaly detection based on historical patterns.

Use **On-Demand** for development, investigation, or one-off validation checks where you don't need continuous alerting.

***

## Related pages

{% content-ref url="/pages/6RrAkEEUumLsBRBdTnw3" %}
[Custom Scheduling For Monitors](/data-quality/monitor-configuration-settings/custom-scheduling-for-monitors)
{% endcontent-ref %}

{% content-ref url="/pages/YpzBR4Fms5X6aZfHSJXj" %}
[Set Up Freshness Monitors](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)
{% endcontent-ref %}

{% content-ref url="/pages/K9nwiEHp0zGa6BN2j0kl" %}
[Set Up Volume Monitors](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)
{% endcontent-ref %}

{% content-ref url="/pages/Dtvx04oJ0XpeTvHOkQQD" %}
[Set Up Field Health Monitors](/data-quality/how-to-set-up-monitors/set-up-field-tests)
{% endcontent-ref %}

{% content-ref url="/pages/PWGhO7SkX4GtJu3jfRBX" %}
[Set Up Custom SQL Monitors](/data-quality/how-to-set-up-monitors/custom-sql-monitors)
{% endcontent-ref %}


# Custom Scheduling For Monitors

The scan frequency options available for Scheduled monitors and the additional settings each frequency exposes.

{% hint style="info" %}
Custom scheduling applies to Scheduled monitors only. On-Demand monitors are triggered manually and do not use a frequency setting.
{% endhint %}

Scan frequency controls how often a Scheduled monitor runs. The frequency you choose also determines how much historical data the ML model collects during training — higher-frequency monitors collect a shorter lookback window; lower-frequency monitors collect a longer one. See the [historical lookback table](/data-quality/anomaly-detection-explained#historical-lookback-by-scan-frequency).

***

## Available frequencies

| Frequency      | Additional settings                               |
| -------------- | ------------------------------------------------- |
| Every 1 hour   | Timezone                                          |
| Every 3 hours  | Timezone                                          |
| Every 6 hours  | Timezone                                          |
| Every 12 hours | Timezone                                          |
| Daily          | Timezone, time of day                             |
| Weekly         | Timezone, day of week, time of day                |
| Monthly        | Timezone, day of month (or last day), time of day |

<figure><img src="/files/Go9hNpX2QHuOH4kp92qu" alt=""><figcaption><p>Frequency selector for a Scheduled monitor</p></figcaption></figure>

***

## Daily scheduling

For daily monitors, you configure:

* **Timezone** — the timezone used to interpret the scheduled time
* **Time of day** — the hour at which the monitor runs (24-hour format)

<figure><img src="/files/LxGRxtdxpfUiUgWjmGxL" alt=""><figcaption><p>Daily scheduling options</p></figcaption></figure>

Schedule daily monitors to run after your ETL jobs complete, so the scan reflects the latest data.

***

## Weekly scheduling

For weekly monitors, you configure:

* **Timezone**
* **Day of week** — which day the monitor runs
* **Time of day**

<figure><img src="/files/0TMRHOVVe9NelJCIcwLN" alt=""><figcaption><p>Weekly scheduling options</p></figcaption></figure>

***

## Monthly scheduling

For monthly monitors, you configure:

* **Timezone**
* **Day of month** — a specific date (up to the 28th) or the last day of the month
* **Time of day**

Use the **last day of the month** option when you need to monitor month-end data regardless of whether the month has 28, 30, or 31 days.

<figure><img src="/files/0Olj8UJauZCj5qkI3ffz" alt=""><figcaption><p>Monthly scheduling options</p></figcaption></figure>

***

## Choosing the right frequency

Match the scan frequency to how often your data actually changes:

* If your pipeline loads data hourly, an hourly monitor catches issues within the same cycle.
* If your pipeline runs once daily (for example, a nightly ETL), a daily monitor scheduled shortly after the load completes is sufficient.
* For reference tables or slowly-changing dimensions, weekly or monthly is appropriate.

{% hint style="warning" %}
Changing scan frequency on a monitor that uses Smart Training triggers a retrain and deletes all historical data. See [Retraining Monitors](/data-quality/monitor-configuration-settings/retraining-monitors).
{% endhint %}

{% content-ref url="/pages/ywQFsioEFNJ8sKDYoSlw" %}
[Available Monitor Modes](/data-quality/monitor-configuration-settings/available-monitor-modes)
{% endcontent-ref %}


# Setting Up Your Data Quality Thresholds

The four threshold types available in Decube monitors, which test types support each, and how thresholds relate to your DQ score.

{% embed url="<https://www.loom.com/share/cd16da42dfc949e48a50d4a537120a2a>" %}

A threshold defines the acceptable range for a monitored metric. When a scan result falls outside the threshold, Decube opens an incident. Choosing the right threshold type for your test determines both what you're measuring and how sensitive the monitor will be.

## The four threshold types

### Absolute

Compares the raw **row count** of failing records against a min/max bound. Use Absolute when you need to reason about a specific number of records — for example, "alert if more than 100 rows have null values."

**Compatible tests:** Null, Unique, Email, UUID, Regex Match

| Setting     | Example value | Effect                                                            |
| ----------- | ------------- | ----------------------------------------------------------------- |
| Max only    | `0`           | Zero failing records permitted — any failure triggers an incident |
| Max only    | `50`          | Up to 50 failing records accepted; incident triggers above that   |
| Min and Max | `10` / `50`   | Incident triggers if failing row count is below 10 or above 50    |

***

### Percentage

Compares the **percentage** of failing records against a min/max bound. Values must be between 0 and 100. Use Percentage when you care about the proportion of bad data relative to the total, rather than the raw count — for example, "alert if more than 2% of emails are invalid."

**Compatible tests:** Null, Unique, Email, UUID, Regex Match

| Setting     | Example value | Effect                                                    |
| ----------- | ------------- | --------------------------------------------------------- |
| Max only    | `0`           | Zero tolerance — any failure triggers an incident         |
| Max only    | `2`           | Up to 2% failures accepted                                |
| Min and Max | `1` / `5`     | Incident triggers if failure rate is below 1% or above 5% |

***

### Positive Range

Compares a **numeric metric value** against a min/max bound, where bound values must be zero or positive. Use Positive Range for statistical metrics that cannot go negative — for example, "alert if the average order value drops below 50 or rises above 500."

**Compatible tests:** Average, Min, Max, String Length

| Setting     | Example value | Effect                                              |
| ----------- | ------------- | --------------------------------------------------- |
| Min only    | `100`         | Alert if the metric drops below 100                 |
| Max only    | `500`         | Alert if the metric exceeds 500                     |
| Min and Max | `100` / `500` | Alert if the metric falls outside the 100–500 range |

***

### Any Range

Compares a **numeric metric value** against a min/max bound, where bound values can be negative. Use Any Range when the monitored column can legitimately hold negative values — for example, a profit/loss column where a value of –1000 to +5000 is expected.

**Compatible tests:** Average, Min, Max

***

### Auto (Smart Training)

When Smart Training is enabled, the ML model learns an expected confidence interval from historical data and sets the effective threshold bounds dynamically. You do not set bounds manually — the model determines what's normal for each scan window.

Auto is only available on Scheduled monitors and requires a timestamp-based row filtering mode.

**Compatible tests:** Null, Unique, Email, UUID, Regex Match, Average, Min, Max, String Length

See [How Anomaly Detection Works](/data-quality/anomaly-detection-explained) for how the model builds and adjusts its confidence interval.

***

## Threshold type compatibility by test

| Test type     | Absolute | Percentage | Positive Range | Any Range | Auto                                 |
| ------------- | -------- | ---------- | -------------- | --------- | ------------------------------------ |
| Null          | ✓        | ✓          | —              | —         | ✓                                    |
| Unique        | ✓        | ✓          | —              | —         | ✓                                    |
| Email         | ✓        | ✓          | —              | —         | ✓                                    |
| UUID          | ✓        | ✓          | —              | —         | ✓                                    |
| Regex Match   | ✓        | ✓          | —              | —         | ✓                                    |
| Average       | —        | —          | ✓              | ✓         | ✓                                    |
| Min           | —        | —          | ✓              | ✓         | ✓                                    |
| Max           | —        | —          | ✓              | ✓         | ✓                                    |
| String Length | —        | —          | ✓              | —         | ✓                                    |
| Cardinality   | —        | —          | —              | —         | Rolling window (no manual threshold) |

***

## Thresholds and your DQ score

Your **DQ score** represents the percentage of records that passed their monitors in a given period. Thresholds determine when a monitor counts a record as failed, so they directly control what your DQ score measures.

The formula is:

<figure><img src="/files/UrNiIGIyOQW2N5ygI1Wx" alt=""><figcaption></figcaption></figure>

**Simple rule of thumb:** your Max threshold is the error budget for your data. If you want a 98% DQ score, set a Max Percentage threshold of 2%.

| Business goal                            | Threshold type | Setting                  | Outcome                                         |
| ---------------------------------------- | -------------- | ------------------------ | ----------------------------------------------- |
| Zero tolerance (e.g., primary keys)      | Absolute       | Max `0`                  | Any null or duplicate triggers an incident      |
| High consistency (e.g., customer emails) | Percentage     | Max `1`                  | Incident triggers when more than 1% are invalid |
| General quality (e.g., optional fields)  | Percentage     | Max `5`                  | Incident triggers when more than 5% are invalid |
| Volume within expected range             | Positive Range | Min `1000` / Max `10000` | Incident if row count falls outside the band    |

### Tuning thresholds over time

Start strict and adjust based on what you observe:

1. **Start at zero or one** for business-critical columns. It's easier to loosen a threshold than to explain why a gap went undetected.
2. **Raise the threshold if false positives accumulate** on non-critical fields — align the setting with what the business genuinely considers a problem.
3. **Review monthly** as data volumes and patterns evolve.

{% content-ref url="/pages/PZ2RI1yYFjLNn2k7GoMG" %}
[Monitor Configuration Reference](/data-quality/monitor-configuration-settings/configuration-reference)
{% endcontent-ref %}


# Retraining Monitors

Which monitor configuration changes trigger a retrain, what data gets deleted when a retrain runs, and how to reconfigure monitors safely.

When you modify certain settings on an existing monitor, Decube discards the monitor's historical data and retrains the ML model from scratch. Understanding which changes trigger this — and what gets deleted — helps you avoid unintended data loss.

{% hint style="danger" %}
A retrain deletes all historical test results, metrics, cases, and incident history for the monitor, including any **open incidents**. Note open incidents before reconfiguring a monitor if you need to preserve that information.
{% endhint %}

## What triggers a retrain

The following configuration changes trigger an automatic retrain:

| Setting changed                                              | Notes                                                                                                                         |
| ------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------- |
| Row creation mode (Timestamp / SQL Expression / All Records) | Changing mode resets the data collection baseline                                                                             |
| Timestamp column                                             | Changing the column changes what the model is trained on                                                                      |
| SQL Expression value                                         | Any change to the expression text                                                                                             |
| Smart Training toggled on                                    | Enabling Smart Training after creation starts a fresh training run                                                            |
| Scan frequency                                               | Changing frequency changes the lookback window (see [How Anomaly Detection Works](/data-quality/anomaly-detection-explained)) |
| Timezone                                                     | Changes the alignment of scan windows                                                                                         |
| Time of day                                                  | Changes when scans run                                                                                                        |
| Day of week                                                  | Applies to weekly-frequency monitors                                                                                          |
| Day of month                                                 | Applies to monthly-frequency monitors                                                                                         |

## What gets deleted during a retrain

When a retrain is triggered, Decube permanently deletes the following for that monitor:

* All historical test results and metric values
* All cases associated with past scan runs
* All incident history
* All **open incidents** — these are closed and deleted, not resolved

There is no recovery path for deleted incident history. If you need to retain a record of open incidents, resolve or export them before making any of the changes listed above.

## What does not trigger a retrain

The following changes are safe to make without triggering a retrain or losing any data:

| Setting                                                            | Safe to change |
| ------------------------------------------------------------------ | -------------- |
| Monitor name                                                       | Yes            |
| Monitor description                                                | Yes            |
| Notification channels                                              | Yes            |
| Incident level                                                     | Yes            |
| Threshold bounds (Absolute, Percentage, Positive Range, Any Range) | Yes            |

## Operational guidance

Before reconfiguring a monitor that uses Smart Training:

1. Go to **All Monitors** and filter to the monitor you plan to change.
2. Review any open incidents and take note of them, or resolve them if they are addressed.
3. Make your configuration change. Decube begins retraining automatically.
4. The monitor enters a training period during which no new incidents are generated. Check the [training duration](/data-quality/anomaly-detection-explained#historical-lookback-by-scan-frequency) for your scan frequency to know when to expect incidents to resume.

{% hint style="info" %}
If you need to change a setting that triggers a retrain on a high-volume production table, consider the timing carefully. The monitor will be inactive during the training period, which can range from hours to days depending on the scan frequency.
{% endhint %}

## Settings that cannot be changed after creation

Some settings are locked after a monitor is created. To change them, you must delete the monitor and create a new one:

* **Asset** (source, schema, dataset)
* **Test type** (for example, you cannot change a Freshness monitor into a Volume monitor)
* **Monitor mode** (Scheduled vs On-Demand)
* **Group by** configuration


# How to set up monitors

Step-by-step setup guides for each monitor type.

1. [Set Up Freshness Monitors](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)
2. [Set Up Volume Monitors](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)
3. [Set Up Field Health Monitors](/data-quality/how-to-set-up-monitors/set-up-field-tests)
4. [Set Up Custom SQL Monitors](/data-quality/how-to-set-up-monitors/custom-sql-monitors)
5. [Grouped-by Monitors](/data-quality/how-to-set-up-monitors/set-up-grouped-by-monitors)
6. [Modify Schema Drift Monitors](/data-quality/how-to-set-up-monitors/set-up-schema-drift-monitors)
7. [Modify Job Failure Monitors (Data Job)](/data-quality/how-to-set-up-monitors/set-up-data-job-job-failure-monitors)
8. [Catalog: Add/Modify Monitor](/data-quality/how-to-set-up-monitors/catalog-add-modify-monitor)


# Catalog: Add/Modify Monitor

This guide walks you through the steps to add and modify monitors from the Asset Details section in the Catalog.

You can add and modify monitors directly from the Data Catalog's Asset Details view, without leaving the context of the asset you're working on.

## Monitor Management Options

### Option 1: Dedicated Config Module (Recommended for Bulk Setup)

* **Best for**: Setting up multiple monitors across different assets
* **Access**: Main navigation Config section
* **Guide**: [Enable Asset Monitoring](/data-quality/enable-asset-monitoring)

### Option 2: Catalog Integration (This Guide)

* **Best for**: Asset-specific monitoring while browsing catalog
* **Access**: Individual asset details pages
* **Context**: Full asset metadata and lineage available

***

## Adding New Monitors from Catalog

<figure><img src="/files/TbtEzUN9cxqpOlnD9fnS" alt=""><figcaption><p>Overview for Add Monitor through Asset Details</p></figcaption></figure>

### Step 2: Create Monitor

1. **Click** the **"Add Monitor"** button
2. **Select** monitor type from the creation form
3. **Configure** settings based on asset characteristics

**Adding a New Monitor**

### Step 3: Choose Monitor Type

Select the appropriate monitor based on your asset needs:

**For Business-Critical Tables:**

<figure><img src="/files/jTwxGwXWnRGXwNGwjmeF" alt=""><figcaption><p>Overview of Modify monitor form in Asset Details</p></figcaption></figure>

**Adding a New Monitor**

* [**Freshness**](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors): Ensure timely data updates
* [**Volume**](/data-quality/how-to-set-up-monitors/set-up-volume-monitors): Validate expected data loads

**For Data Quality Validation:**

* [**Field Health**](/data-quality/how-to-set-up-monitors/set-up-field-tests): Column-level quality checks
* [**Custom SQL**](/data-quality/how-to-set-up-monitors/custom-sql-monitors): Business rule validation

**Automatically Enabled:**

* **Schema Drift**: Already active for structural change detection

***

## Modifying Existing Monitors

### Quick Modification Process

1. **Locate** the monitor in the Monitor tab
2. **Click** the **ellipsis menu (︙)** next to the target monitor
3. **Select** "View Monitor" to access modification options

<figure><img src="/files/yPSk7eni5zZ4FWBWQZRa" alt=""><figcaption><p>Monitor modification interface with edit, disable, and delete options</p></figcaption></figure>

### Available Actions

> ⚠️ **Hint:** Modifying the frequency or row creation settings of auto-thresholded monitors may trigger retraining. This will remove all historical values from the monitor.

**✏️ Modify Settings:**

* Update thresholds and sensitivity
* Change monitoring frequency
* Adjust alert configurations

**⏸️ Enable/Disable:**

* Temporarily pause monitoring without losing configuration
* Useful during maintenance or data migration periods

**🗑️ Delete Monitor:**

* Permanently remove monitor and historical data
* Use with caution - action cannot be undone

<figure><img src="/files/nJNNH3AlI3yb4kFDX6gQ" alt=""><figcaption><p>Overview of Create a new monitor form</p></figcaption></figure>


# Set Up Freshness Monitors

Set up Freshness monitors to detect when data stops arriving on its expected schedule.

A Freshness monitor learns when your data normally arrives and alerts you when it doesn't show up as expected. Unlike a simple staleness check, the scheduled Freshness monitor builds a probability model from historical arrival patterns — so it knows not to alert on weekends if your data never arrives on weekends.

## Freshness vs Volume: which to use

| Goal                                              | Use           |
| ------------------------------------------------- | ------------- |
| Detect that data **arrived** (or didn't)          | **Freshness** |
| Detect that the **right amount** of data arrived  | **Volume**    |
| Table receives data on a known schedule           | **Freshness** |
| Table grows by a predictable row count per period | **Volume**    |

If you need both signals, create one monitor of each type on the same table.

## How the scheduled Freshness monitor works

The scheduled Freshness monitor uses an ML model trained on your table's historical write timestamps. The model learns the expected arrival windows for your data and flags a run as anomalous when the probability of data having arrived drops below 50%.

This means:

* If your pipeline never runs on weekends, the model accounts for this and does not raise weekend incidents.
* If your data normally arrives between 08:00 and 10:00, a scan at 07:00 that finds no new data will not trigger an alert — the model knows it is too early.

Each incident for a Freshness monitor includes the **`time_since_last_write`** metric, which shows exactly how long ago the table last received a write. Use this to triage whether a delay is minor or critical.

## How the on-demand Freshness monitor works

The on-demand Freshness monitor performs a simpler presence check: it queries whether any new rows exist since the lookback period you specify at run time. It does not use the ML model or learn arrival patterns — it returns a pass or fail based purely on whether rows are present.

## Before you begin

* You need at least one data source connected and a table available under that source.
* To use Smart Training, the table must have a timestamp column (or you must provide an SQL expression that produces one). Smart Training is not available in On-Demand mode.

***

## Step 1: Set up

1. In the **Data Quality** module, go to the **Config** tab and select **Create**.
2. Select the **Freshness** monitor card.
3. In the **Create a New Monitor** form, select your **Source** (Schema is optional) and **Dataset**.
4. Choose **Monitor mode**: **Scheduled** or **On-Demand**.
5. Optionally enable **Grouped By** — select a column to group on and click **Validate** to confirm the column is valid.
6. Click **Proceed to Monitor Setup**.

{% hint style="info" %}
Grouped By creates one sub-monitor per distinct value in the group column. The [Grouped-By Monitors](/data-quality/how-to-set-up-monitors/set-up-grouped-by-monitors) page explains the 100-distinct-values limit and other constraints.
{% endhint %}

***

## Step 2: Configure — Scheduled monitor

Complete the required fields in the **Configure** form:

| Field                   | Description                                                                                                                                                                                   |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Monitor Name**        | A descriptive name for this monitor. You can create multiple Freshness monitors on the same table.                                                                                            |
| **Monitor Description** | Optional.                                                                                                                                                                                     |
| **Row Creation**        | How Decube identifies new rows: **Timestamp** (select a timestamp column), **SQL Expression** (provide an expression that produces a timestamp), or **All Records**.                          |
| **Smart Training**      | Toggle on to train the model on historical data. Requires Timestamp or SQL Expression row creation. Enabling Smart Training also makes the **Lookback Period** selectable.                    |
| **Frequency**           | How often the monitor scans. See [Custom Scheduling for Monitors](https://github.com/DecubeIO/decube-docs/blob/public/data-quality/how-to-set-up-monitors/custom-scheduling-for-monitors.md). |
| **Incident Level**      | Severity assigned to incidents this monitor opens.                                                                                                                                            |

{% hint style="warning" %}
**All Records** mode does not support Smart Training. The monitor runs without a probability model and compares total row counts directly.
{% endhint %}

### SQL Expression

Use an SQL Expression when your table stores timestamps in a non-standard format (string, Unix epoch, or split date/time columns). The expression must produce a valid timestamp in your data source's SQL dialect.

| Format                       | BigQuery example                                           | PostgreSQL example                             |
| ---------------------------- | ---------------------------------------------------------- | ---------------------------------------------- |
| String → timestamp           | `CAST(your_col AS DATETIME)`                               | `your_col::timestamp`                          |
| Unix seconds → timestamp     | `TIMESTAMP_SECONDS(your_col)`                              | `TO_TIMESTAMP(your_col)`                       |
| Separate date + time columns | `PARSE_DATETIME('%F %T', CONCAT(date_col, ' ', time_col))` | `(date_col \|\| ' ' \|\| time_col)::timestamp` |

Validating the expression before saving is required.

### Notifications

Turn on **Notify default channel** to route incidents to a specific email or Slack channel. Click **Submit** to create the monitor.

***

## Step 2: Configure — On-Demand monitor

On-Demand monitors do not use Smart Training, Auto Threshold, frequency scheduling, or Grouped By.

| Field                   | Description                                                                                |
| ----------------------- | ------------------------------------------------------------------------------------------ |
| **Monitor Name**        | A descriptive name.                                                                        |
| **Monitor Description** | Optional.                                                                                  |
| **Row Creation**        | **Timestamp** or **SQL Expression** only — All Records is not available in On-Demand mode. |
| **Lookback Period**     | The time window to check for new rows when the monitor runs.                               |
| **Incident Level**      | Severity assigned to incidents this monitor opens.                                         |

To finish:

* Click **Save** to create the monitor without running it immediately.
* Click **Save and Run** to create and run the monitor straight away.

After creation, you can run the monitor again from **All Monitors** by clicking the ellipsis (︙) and selecting **View Monitor**, then **Run once**.

***

## Related pages

{% content-ref url="/pages/K9nwiEHp0zGa6BN2j0kl" %}
[Set Up Volume Monitors](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)
{% endcontent-ref %}

{% content-ref url="/pages/12tZSWlWRX4XEXBTD4L6" %}
[How Anomaly Detection Works](/data-quality/anomaly-detection-explained)
{% endcontent-ref %}

{% content-ref url="/pages/KgXixMx9V2joZYvvAjL9" %}
[Retraining Monitors](/data-quality/monitor-configuration-settings/retraining-monitors)
{% endcontent-ref %}


# Set Up Volume Monitors

Set up Volume monitors to detect when the number of new rows arriving in a table is abnormally high or low.

A Volume monitor tracks how many rows are added to a table in each scan window and alerts you when that increment falls outside the expected range. It is the right tool when you care about *how much* data arrived — not just whether it arrived.

## Freshness vs Volume: which to use

| Goal                                              | Use           |
| ------------------------------------------------- | ------------- |
| Detect that data **arrived** (or didn't)          | **Freshness** |
| Detect that the **right amount** of data arrived  | **Volume**    |
| Table receives data on a known schedule           | **Freshness** |
| Table grows by a predictable row count per period | **Volume**    |

If you need both signals, create one monitor of each type on the same table.

## How the Volume monitor works

The Volume monitor measures the **increment** — the number of new rows added since the previous scan — not the total row count of the table. The ML model trains on historical increments to establish a normal range, and raises an incident when a new increment falls outside it.

{% hint style="warning" %}
**Volume and row deletions**: if rows are deleted from a table and you use **All Records** mode, the measured increment can be negative or artificially small. This produces incorrect deltas. Use **Timestamp** row creation on tables where rows may be deleted.
{% endhint %}

### Backfill warning

If a large historical backfill is loaded into a monitored table, the increment for that scan period will be far above normal. The model will flag this as an anomaly. If you plan a backfill, either mute the monitor during the load or acknowledge the resulting incident as expected behaviour.

### Group by not supported

Volume monitors do not support Group By. If you need volume tracking broken down by a column value, use a Custom SQL monitor with a `GROUP BY` clause.

### Low-volume tables

On tables with very few rows per period, the Volume monitor may hit the [sparse-data threshold](/data-quality/anomaly-detection-explained#sparse-data-and-silent-skipping) (fewer than 5 valid points in 30 observations) and silently skip scans. For low-volume tables, a manual threshold (Absolute or Percentage) is more reliable than Smart Training.

***

## Before you begin

* You need at least one data source connected and a table available under that source.
* To use Smart Training, the table must have a timestamp column (or you must provide an SQL expression that produces one).

***

## Step 1: Set up

1. In the **Data Quality** module, go to the **Config** tab and select **Create**.
2. Select the **Volume** monitor card.
3. In the **Create a New Monitor** form, select your **Source** (Schema is optional) and **Dataset**.
4. Choose **Monitor mode**: **Scheduled** or **On-Demand**.
5. Click **Proceed to Monitor Setup**.

***

## Step 2: Configure — Scheduled monitor

| Field                   | Description                                                                                                                                                                                   |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Monitor Name**        | A descriptive name for this monitor.                                                                                                                                                          |
| **Monitor Description** | Optional.                                                                                                                                                                                     |
| **Row Creation**        | How Decube identifies new rows: **Timestamp** (select a timestamp column), **SQL Expression** (provide an expression that produces a timestamp), or **All Records**.                          |
| **Smart Training**      | Toggle on to train the model on historical increments. Requires Timestamp or SQL Expression.                                                                                                  |
| **Frequency**           | How often the monitor scans. See [Custom Scheduling for Monitors](https://github.com/DecubeIO/decube-docs/blob/public/data-quality/how-to-set-up-monitors/custom-scheduling-for-monitors.md). |
| **Incident Level**      | Severity assigned to incidents this monitor opens.                                                                                                                                            |

{% hint style="warning" %}
**All Records** mode does not support Smart Training and can produce incorrect deltas if rows are deleted. Use Timestamp or SQL Expression row creation whenever possible.
{% endhint %}

### SQL Expression

Use an SQL Expression when your table stores timestamps in a non-standard format. See the [SQL Expression examples on the Freshness setup page](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors#sql-expression) for common conversion patterns. Validating the expression before saving is required.

### Notifications

Turn on **Notify default channel** to route incidents to a specific email or Slack channel. Click **Submit** to create the monitor.

***

## Step 2: Configure — On-Demand monitor

On-Demand monitors do not use Smart Training, Auto Threshold, or frequency scheduling.

| Field                   | Description                                        |
| ----------------------- | -------------------------------------------------- |
| **Monitor Name**        | A descriptive name.                                |
| **Monitor Description** | Optional.                                          |
| **Row Creation**        | **Timestamp** or **SQL Expression** only.          |
| **Lookback Period**     | The time window to check when the monitor runs.    |
| **Incident Level**      | Severity assigned to incidents this monitor opens. |

To finish:

* Click **Save** to create the monitor without running it immediately.
* Click **Save and Run** to create and run the monitor straight away.

After creation, you can run the monitor again from **All Monitors** by clicking the ellipsis (︙) and selecting **View Monitor**, then **Run once**.

***

## Related pages

{% content-ref url="/pages/YpzBR4Fms5X6aZfHSJXj" %}
[Set Up Freshness Monitors](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)
{% endcontent-ref %}

{% content-ref url="/pages/12tZSWlWRX4XEXBTD4L6" %}
[How Anomaly Detection Works](/data-quality/anomaly-detection-explained)
{% endcontent-ref %}

{% content-ref url="/pages/KgXixMx9V2joZYvvAjL9" %}
[Retraining Monitors](/data-quality/monitor-configuration-settings/retraining-monitors)
{% endcontent-ref %}


# Set Up Custom SQL Monitors

Write SQL to define custom validation logic and monitor it on a schedule or on demand.

Custom SQL monitors let you define validation logic that the built-in test types cannot express — cross-table checks, complex business rules, aggregation-based assertions. You write a SQL query that returns the failing rows; Decube wraps it in a `SELECT COUNT(*) FROM (<your_sql>) AS base` and opens an incident when the row count is greater than zero.

## How Custom SQL works

**Custom SQL attaches to a Source, not a table.** In the setup form you select a data source connection (BigQuery project, Snowflake account, Postgres database, etc.) rather than a specific table. Your SQL query itself determines which tables are scanned.

**Your query must return rows, not a scalar count.** Because Decube wraps your SQL in an outer `COUNT(*)`, writing `SELECT COUNT(*) FROM table WHERE ...` is a common mistake. If there are no failing records, that query returns one row containing the value `0` — and the outer wrapper counts that as 1 row, triggering an incident on every scan. Write your SQL to return the actual failing records instead:

```sql
-- ✓ Correct — returns failing rows
SELECT order_id, amount
FROM orders
WHERE amount < 0

-- ✗ Incorrect — returns a scalar; will always trigger
SELECT COUNT(*) FROM orders WHERE amount < 0
```

**Custom SQL never uses Smart Training.** The Auto threshold mode and the ML confidence interval are not available for Custom SQL monitors. All thresholds must be set manually (the default is `row_count > 0`).

**Changing the SQL query does not trigger a retrain.** Since Custom SQL does not use Smart Training, modifying the SQL expression only requires re-validation — it does not delete historical data.

**Query timeout is 55 seconds.** If your query does not return within 55 seconds, the scan fails with a timeout error. For complex queries on large tables, optimise with filters and avoid full-table scans.

***

## Before you begin

* You need a connected data source with appropriate read permissions on the tables your SQL will reference.
* Your SQL must be written in the dialect of the selected data source (BigQuery standard SQL, Snowflake SQL, PostgreSQL, etc.).
* Validate your query in the form before saving — validation is required.

***

## Step 1: Set up

1. In the **Data Quality** module, go to the **Config** tab and select **Create**.
2. Select the **Custom SQL** monitor card.
3. Select a **Data Source**.
4. Choose **Monitor mode**: Scheduled or On-Demand.
5. Optionally enable **Grouped By** — see [Custom SQL Group By](/data-quality/how-to-set-up-monitors/set-up-grouped-by-monitors#configure-on-demand-schedule-monitor-grouped-by-for-custom-sql) for how group mapping works with Custom SQL.
6. Click **Proceed to Monitor Setup**.

***

## Step 2: Configure — Scheduled monitor

| Field                   | Description                                                                                                                                                       |
| ----------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Monitor Name**        | A descriptive name. You can create multiple Custom SQL monitors on the same source.                                                                               |
| **Monitor Description** | Optional.                                                                                                                                                         |
| **Frequency**           | How often the monitor runs. See [Custom Scheduling for Monitors](/data-quality/monitor-configuration-settings/custom-scheduling-for-monitors).                    |
| **Custom SQL query**    | Enter your SQL, then click **Validate**. A "Query successfully validated" message confirms it. If you edit the query after validating, re-validate before saving. |
| **Incident Level**      | Severity assigned to incidents this monitor opens.                                                                                                                |

Turn on **Notify default channel** to route incidents to an email address or Slack channel, then click **Submit**.

***

## Step 2: Configure — On-Demand monitor

On-Demand monitors do not use frequency scheduling.

| Field                   | Description                                        |
| ----------------------- | -------------------------------------------------- |
| **Monitor Name**        | A descriptive name.                                |
| **Monitor Description** | Optional.                                          |
| **Custom SQL query**    | Enter your SQL and click **Validate**.             |
| **Incident Level**      | Severity assigned to incidents this monitor opens. |

To finish:

* Click **Save** to create the monitor without running it immediately.
* Click **Save and Run** to create and run the monitor straight away.

To run the monitor again, go to **All Monitors**, click the ellipsis (︙) on the monitor, select **View Monitor**, then **Run once**.

***

## Including Custom SQL in the DQ Scorecard <a href="#dq-scorecard" id="dq-scorecard"></a>

{% hint style="info" %}
By default, Custom SQL monitors are not included in the [DQ Scorecard](/dashboard/health-score) or [DQ Reports](/reports/asset-report-data-quality-scorecard). The scorecard requires both an error row count and a total row count to calculate a percentage score. Custom SQL monitors only produce a row count of failures — the total row count must be provided separately.
{% endhint %}

To include a Custom SQL monitor in the DQ score:

1. In the monitor configuration, enable **Include monitor into Data Quality Scorecard**.
2. Provide a **Total Row Count Query** — a SQL query that returns the total number of rows in the search space your main query checks against.

Both queries run on every scan. The score for that monitor is calculated as:

```
DQ Score = 1 - (error_row_count / total_row_count)
```

{% embed url="<https://www.loom.com/share/fd585bcd7602408c965a2a6a4cbf8a2a?sid=4d031395-953c-4caf-8c46-1f4e41bed154>" %}

### Example

**Scenario:** Detect rows with invalid `event_type` values in `analytics.events` and include the monitor in the DQ Scorecard.

{% code title="Custom SQL query (error rows)" overflow="wrap" %}

```sql
SELECT COUNT(*) AS value
FROM analytics.events
WHERE event_type NOT IN ('click', 'view', 'purchase')
  AND event_time >= '2025-09-01'::timestamp
  AND event_time < '2025-09-02'::timestamp
```

{% endcode %}

{% code title="Total Row Count Query" overflow="wrap" %}

```sql
SELECT COUNT(*) AS value
FROM analytics.events
WHERE event_time >= '2025-09-01'::timestamp
  AND event_time < '2025-09-02'::timestamp
```

{% endcode %}

{% hint style="warning" %}
If `total_row_count` is ever less than `error_row_count` (for example, due to a logic error in the total query), that observation is treated as invalid and omitted from the DQ score calculation. Take care to ensure both queries operate over the same scope.
{% endhint %}

### Viewing scan history with row counts

When a Total Row Count Query is enabled, the monitor's history in **Config > All Monitors > Monitor Info** shows an additional `row_count` column alongside the error count for each scan.

***

## Tips and Tricks

### Label invalidation

Alert when records within a recent window have a specific status value:

```sql
SELECT *
FROM public.sale
WHERE created_at BETWEEN NOW() - '30mins'::INTERVAL AND NOW()
  AND status IN ('pending', 'rejected')
```

### Value sum exceeds threshold

Alert when the sum of a value over a time window crosses a threshold:

```sql
SELECT 1
FROM public.sale
WHERE created_at BETWEEN NOW() - '30mins'::INTERVAL AND NOW()
HAVING SUM(price * quantity) > 10000
```

### Cross-table discrepancy

Alert when two tables fall out of sync:

```sql
WITH sale AS (
    SELECT SUM(price * quantity) AS amount
    FROM public.sale
    WHERE created_at BETWEEN NOW() - '30mins'::INTERVAL AND NOW()
),
in_flow AS (
    SELECT SUM(amount) AS amount
    FROM public.in_flow
    WHERE created_at BETWEEN NOW() - '30mins'::INTERVAL AND NOW()
)
SELECT (SELECT amount FROM sale) - (SELECT amount FROM in_flow) AS discrepancy
FROM (SELECT 1) AS dummy
WHERE (SELECT amount FROM sale) - (SELECT amount FROM in_flow) > 10000
```

***

{% content-ref url="/pages/PZ2RI1yYFjLNn2k7GoMG" %}
[Monitor Configuration Reference](/data-quality/monitor-configuration-settings/configuration-reference)
{% endcontent-ref %}


# Set Up Field Health Monitors

Follow these steps to configure monitors for specific field tests, available as On-Demand and Scheduled modes.

{% hint style="info" %}
On-Demand Monitors are **not applicable** for the "Cardinality" test type.
{% endhint %}

**Enabling Field Health Monitoring**

To enable field health monitoring:

1. Navigate to the **Config** landing page.
2. Select the **“Field Health”** card.

<figure><img src="/files/5SAy4u88rce2PXqVvj7Q" alt=""><figcaption><p>Selecting Field Health Card</p></figcaption></figure>

You can also activate field health monitoring within the "Asset Details" section of the Data Catalog Module. This can be done via the "Monitors" tab.

For more detailed information you can refer to below link

{% content-ref url="/pages/eRwubVBjQKYfSlbVC3T1" %}
[Catalog: Add/Modify Monitor](/data-quality/how-to-set-up-monitors/catalog-add-modify-monitor)
{% endcontent-ref %}

Once selected, you’ll be redirected to the **“Create a New Monitor”** form.

* The **“Create a New Monitor”** form consists of two steps:

  **Setup**

  **Configure**

{% hint style="info" %}
The form fields will become available as you select the mandatory options.
{% endhint %}

### Step 1: Set-up

* Choose the `test type` from the dropdown. The available test types are:
  1. **Null** — Monitors null values. Supports Absolute (row count), Percentage, and Auto threshold modes.
  2. **Unique** — Monitors duplicate values. Supports Absolute, Percentage, and Auto threshold modes.
  3. **Average** — Computes the column average and compares it against a Range threshold or Auto mode. Numeric columns only.
  4. **Min** — Identifies the minimum value in a column. Supports Range and Auto threshold modes. Numeric columns only.
  5. **Max** — Identifies the maximum value in a column. Supports Range and Auto threshold modes. Numeric columns only.
  6. **Cardinality** — Tracks the number of distinct values in a column using a 5-period rolling window. See [How Cardinality detection works](#cardinality-detection) below. Scheduled mode only.
  7. **String Length** — Validates string lengths. Supports Range and Auto threshold modes.
  8. **Email** — Monitors invalid email addresses. Supports Absolute, Percentage, and Auto threshold modes.
  9. **UUID** — Monitors invalid UUIDs. Supports Absolute, Percentage, and Auto threshold modes.
  10. **Regex Match** — Validates values against a user-defined pattern. Specify the pattern separately from the threshold. Supports Absolute, Percentage, and Auto threshold modes.
* **Important Note for Synapse/SQL Server Users:**

When setting up the Match REGEX threshold for data quality monitoring:

• `Synapse/SQL Server` requires `string` matching patterns instead of standard REGEX syntax.

• Ensure that the value entered in the “Set Threshold” field aligns with the string pattern supported by `Synapse/SQL Server`.

{% hint style="info" %}
If the current preset field tests is not sufficient to run the specific test you require, you can also create a test via custom SQL script.

[Custom SQL Monitor](/data-quality/how-to-set-up-monitors/custom-sql-monitors)
{% endhint %}

* Select the `data source` **&** Filter by `Schema` is **optional.**
* Select the `dataset` and `Column` you want to monitor.
* Choose `Monitor mode`: Scheduled or On-Demand

{% hint style="info" %}
For a detailed understanding of monitor modes, check out [**Available Monitor Modes**](/data-quality/monitor-configuration-settings/available-monitor-modes)
{% endhint %}

<figure><img src="/files/Hv3WygzDL8cfJ5sI8KFN" alt=""><figcaption></figcaption></figure>

* **Enable “Grouped By” (if applicable)** by toggling the switch.
* Select the column for grouping and click **Validate**.
* A success message (**“Column is valid to be grouped by”**) confirms validation.

<figure><img src="/files/66IICwpw2FGASU27FYSb" alt=""><figcaption><p>Grouped-By Disabled</p></figcaption></figure>

<figure><img src="/files/9hmK7NX6RGGJppwtySd0" alt=""><figcaption><p>Grouped-by enabled</p></figcaption></figure>

* Click **“Proceed to Monitor Setup”** to move to the next step.

<figure><img src="/files/EasaTvDLH7EYBzVRV6tQ" alt=""><figcaption><p>Choose Monitor Mode</p></figcaption></figure>

### Step 2: Configure: Scheduled Monitor

* Once you proceed to setup, you’ll reach the **”Configure”** page, where you can review your previous selections.

Complete the required fields in the **Configure** form:

{% hint style="info" %}

* Can create multiple tests for each test type per column/table

* Able to add a test name to differentiate monitors created
  {% endhint %}

* **Monitor Name**

* **Monitor Description** is Optional.

* **Row Creation:** Select the Row creation from given options:
  * **Timestamp** (Select Timestamp from the dropdown column)
  * **Validation for SQL Expression** (when SQL Expression is chosen)
  * **All Records**
  * **Enable Smart Training(Optional):** Train your monitor on historical data to reduce the training period

* **Frequency** (Learn more about [**Custom Scheduling For Monitors**](/data-quality/monitor-configuration-settings/custom-scheduling-for-monitors))

* **Threshold Mode:** Select the appropriate threshold type based on your test:
  * **Absolute (Row Count):** Set thresholds based on number of rows (e.g., "fail when > 100 invalid rows")
    * Available for: Null, Unique, Email, UUID, Regex Match
    * Input: Non-negative integer values for min/max bounds
  * **Percentage:** Set thresholds based on percentage (e.g., "fail when > 5% nulls")
    * Available for: Null, Unique, Email, UUID, Regex Match
    * Input: Integer values between 0-100 for min/max bounds
  * **Auto (Machine Learning):** Let the system learn acceptable patterns from historical data
    * Only available for scheduled monitors
    * Available for: Null, Unique, Email, UUID, Regex Match, Average, Min, Max, String Length
    * No manual threshold input required
  * **Range:** Set numeric range thresholds for statistical tests
    * Available for: Average, Min, Max, String Length
    * Input: Min/max bounds (at least one required)

* **Set Threshold:** Configure min/max values based on selected threshold mode
  * At least one bound (min or max) is required
  * For percentage mode: values must be 0-100
  * For absolute mode: values must be non-negative integers
  * For range mode: constraints depend on test type

* **Quality Dimension (Optional) for more understanding refer to** [supported test types](#supported-monitors-with-default-quality-dimension)

{% hint style="info" %}
When using **SQL Expression**, validating your query is compulsory. Ensure that your query is written in the dialect compatible with your linked data source, as illustrated below:

**Google BigQuery** - `CAST(your_string_column AS DATETIME)`

**PostgreSQL** - `your_timestamp_column::timestamp`

When working with Google BigQuery, you can review the provided documentation for further details here.
{% endhint %}

{% hint style="info" %}
**Smart Training** requires Row Creation to be selected.

To activate Smart Training in Row Creation:

* Users should initially select the timestamp.
* If SQL Expression is chosen for row creation, validate the SQL Expression before saving.
  {% endhint %}

{% hint style="info" %}
**Regex Match Test:** When configuring a Regex Match test, you'll specify the regex pattern in a separate field in the test configuration (not in the threshold). The threshold controls how many non-matching values will trigger an incident.
{% endhint %}

### Cardinality detection

Cardinality uses a rolling-window approach rather than the ML model used by other Auto tests. Each scan compares the current distinct-value count against the average of the **preceding 5 scan results**. If the current count falls outside the expected range based on that window, Decube opens an incident.

Because this method derives its baseline from recent history rather than a training run, Cardinality does not have a training period and is not affected by the [sparse-data threshold](/data-quality/anomaly-detection-explained#sparse-data-and-silent-skipping) that applies to Smart Training monitors. The rolling window does require at least 5 prior scan results before it can produce meaningful comparisons — a newly created Cardinality monitor will not raise incidents until 5 scans have completed.

{% hint style="info" %}
Cardinality is only supported in **Scheduled** mode. On-Demand mode is not available for this test type.
{% endhint %}

***

<figure><img src="/files/QCQmRZbczx0QDzDOtcDC" alt=""><figcaption><p>Overview for setting-up frequency</p></figcaption></figure>

<figure><img src="/files/4ZXlOKTUiVoBLgV55ezh" alt=""><figcaption><p>Overview of set-up monitor with supported quality dimension</p></figcaption></figure>

### [Get Notified/Custom Alert](#get-notified-custom-alert)

* To set custom alerts, turn on the `"Notify default channel"` toggle, then select your preferred alert channels (email or Slack) from the dropdown.
  * You can select the desired alert channels in the dropdown.
  * Mention the address or channel name in the field.
* At last Specify the **Incident Level**.

<figure><img src="/files/nAKTlcqfooVEsnKo9sEz" alt=""><figcaption><p>Setting-up notification/custom alert</p></figcaption></figure>

* Click on `Submit` and your monitor is created successfully.
* Once monitor is created successfully you will be redirected to ALL MONITORS tab.

### Step 2: Configure: On-demand Monitor

{% hint style="info" %}
**Note: Key Differences:** i. The "`Frequency`" field is not relevant for configuring any On Demand monitors and is therefore neglected.

ii. “`Enable Smart Training`”"`Auto Threshold`" options is neglected when it comes to setting up any On Demand Monitor.

iii. `Grouped By` is not available for On-demand monitor mode.
{% endhint %}

* Select **On-Demand** as the Monitor Mode and click **“Proceed to Monitor Setup”**.

<figure><img src="/files/KrA946COD5RejnZzIxTJ" alt=""><figcaption><p>Choosing monitor mode</p></figcaption></figure>

* Complete the required fields in the **Configure** form:

{% hint style="info" %}

* Can create multiple tests for each test type per column/table
* Able to add a test name to differentiate monitors created
  {% endhint %}

- **Monitor Name**
- **Monitor Description** is Optional.
- **Row Creation:** Select the Row creation from given options:
  * **Timestamp** (Select a timestamp column from the dropdown)
  * **Validation for SQL Expression** (when SQL Expression is chosen)
  * **All Records**
- **Lookback Period**
- **Threshold Mode:** Select the appropriate threshold type:
  * **Absolute (Row Count):** For Null, Unique, Email, UUID, Regex Match
  * **Percentage:** For Null, Unique, Email, UUID, Regex Match
  * **Range:** For Average, Min, Max, String Length
  * Note: Auto threshold is not available for on-demand monitors
- **Set Threshold:** Configure min/max values (at least one required)
- **Quality Dimension (Optional) for more understanding refer to** [supported test types](#supported-monitors-with-default-quality-dimension)
- **Incident Levels**

<figure><img src="/files/7U50oTzV6KbvRz8LnNkj" alt=""><figcaption><p>On-demand monitor configuration</p></figcaption></figure>

{% hint style="info" %}
**Custom Notifications:** Custom alerts can be configured as in scheduled monitors.

[**Get Notified/Custom Alert**](#get-notified-custom-alert)
{% endhint %}

**Finalizing On-Demand Monitor Setup**

* To finish, choose one of the following:
  * `Save` — creates the monitor without running it immediately.
  * `Save and Run` — creates and runs the monitor straight away. To run it again later, navigate to **All Monitors**.
  * After selecting the above option you will be redirected to ALL MONITORS tab.
* **Modify Monitoring**

To modify an existing monitor:

1. Go to [**All Monitors**.](/data-quality/enable-asset-monitoring#all-monitors-tab)
2. Click the ellipsis (︙) and select **View Monitor**.
3. Click on `Run once` to run the monitor manually.

<figure><img src="/files/w6ogeOB3ZlK0W9n7Webl" alt=""><figcaption><p>Modify Monitor from All Monitors</p></figcaption></figure>


# Grouped-by Monitors

This document provides conceptual information for leveraging the Grouped By option in your data quality workflows.

By utilizing Group By monitors, data teams have the ability to highlight and define specific segments within a table, like those aggregated by values in a dimension column. Following this segmentation, monitors can be applied. This functionality enables teams to not only track the overall row count of a table but also monitor counts of its individual logical subdivisions.

The configuration process for Grouped-by monitors differs from the setup routine of other monitor types, here's how you begin setting up for Grouped-by monitors.

### Setting Up Grouped-by monitoring

For initial steps you can refer to below links:

* [Set Up Freshness Monitors](/data-quality/how-to-set-up-monitors/set-up-freshness-monitors)
* [Set Up Volume Monitors](/data-quality/how-to-set-up-monitors/set-up-volume-monitors)
* [Set Up Custom SQL Monitors](/data-quality/how-to-set-up-monitors/custom-sql-monitors)
* [Set Up Field Health monitors](/data-quality/how-to-set-up-monitors/set-up-field-tests)

To configure Grouped-by monitors in Field Health follow below steps:

* **Enable “Grouped By” (if applicable)** by toggling the switch.
  * Select the column for grouping and click **Validate**.
  * A success message (**“Column is valid to be grouped by”**) confirms validation.

<figure><img src="/files/66IICwpw2FGASU27FYSb" alt=""><figcaption><p>Grouped-by diabled</p></figcaption></figure>

<figure><img src="/files/9hmK7NX6RGGJppwtySd0" alt=""><figcaption><p>Grouped-by enabled</p></figcaption></figure>

* Click **“Proceed to Monitor Setup”** to move to the next step.

<figure><img src="/files/EasaTvDLH7EYBzVRV6tQ" alt=""><figcaption><p>Setting-up monitor grouped-by disabled</p></figcaption></figure>

### Configure: On-demand /Schedule Monitor Grouped By for Custom SQL

* Once you select On-demand/Schedule as monitor mode, next step is click on `Proceed to monitor setup` button.

<figure><img src="/files/ygOaNTJIkraQS9uEeKEd" alt=""><figcaption><p>Grouped-by enabled</p></figcaption></figure>

* Once you click on `Proceed to monitor setup` button, you will reach the next page (Configure).

<figure><img src="/files/qfkWFZiN7Yk4wzDV1TVJ" alt=""><figcaption><p>Overview for SQL Query validation</p></figcaption></figure>

* **Result Set Group by column:** Select from dropdown, This is the column from your query result above that contains the value to group by.
* **Result Set Mapping:** Choose the table and column where your distinct value are pulled from.

***

### Configure Grouped-by (Other Tests)

{% hint style="info" %}
Note that if the distinct values in your selected column exceeds more than 100, you will not be able to add that as a group-by column. Please reach out to us if support is required on this.
{% endhint %}

* **Fetch Values:** Click on Fetch Values button and Select group by fetching values from the grouped-by column selected
* **Search and select multiple columns:** Select from dropdown

<figure><img src="/files/cPjj3aL6gjR47V00Tf31" alt=""><figcaption><p>Grouped-by Configuration</p></figcaption></figure>


# Modify Schema Drift Monitors

Here's how you set up a Schema Drift monitor.

Schema Drift Monitors are automatically activated for all tables upon connecting a data source. They identify alterations in the schema, such as the addition, deletion, or modification of tables or columns, and data type changes. Such adjustments can potentially lead to compatibility challenges with the corresponding database and the applications leveraging these tables.

{% hint style="info" %}
Schema Drift monitors do not need to be created, however the settings can be modified.
{% endhint %}

* Under **All Monitors tab** , click on the **Schema Drift** pill to view the list of Schema Drift monitors.
* Click on the **ellipsis menu (⋮)** to view the monitor. The **Modify Monitor** modal will pop up and gives you option to modify, disable and delete monitor.

<figure><img src="/files/hXBjNczap15P1dJ9QJUn" alt=""><figcaption><p>Pill selection for schema</p></figcaption></figure>

<figure><img src="/files/sHf1oMoQ2dOR1pRvAjhB" alt=""><figcaption><p>Modify schema drift through All Monitors</p></figcaption></figure>


# Modify Job Failure Monitors (Data Job)

Here's how you manage monitors for Job Failure (Data Jobs).

{% hint style="info" %}
**Job Failure Monitors** for Data Jobs are created automatically by Decube upon connecting ETL-type sources that has Data Jobs.
{% endhint %}

> **What is Data Job/Job Failure Monitoring?**
>
> It tracks **failed jobs** to identify processing issues, ensuring data pipelines run smoothly:
>
> * Job Failure does not require **manual creation** like other monitors. Instead, it is configured and adjusted separately.
> * This avoids confusion about when and how these monitors should be used.

To effectively detect anomalies in your data transformation, our system offers specialized Data Job/Job Failure monitors. Here's a step-by-step guide to get you started:

1. **Navigate to the Config Page, All Monitors tab:** This is where all monitor types options are available under Test Name column.
2. **Locate the "Job Failure/Data Job" Pill:** Among the various pills, you'll find a pill labeled "Data Job" under the "All Monitors" tab. You can also apply filter for Job failure test name.
3. **Select the "Job Failure/Data Job" Pill:** By selecting this, you'll initiate the modifying monitor setup process for your Data Job/Job Failure monitor

<figure><img src="/files/fVnxgDjLo8fRXRmVZ4ZP" alt=""><figcaption><p>All Monitors tab and Data Job pill</p></figcaption></figure>

<figure><img src="/files/cZZGnGNz2sg0tVkuCXEJ" alt=""><figcaption><p>Data Job/Job Failure Monitors</p></figcaption></figure>

### Data Job/Job Failure Monitoring Setup

Click the ellipsis button ( : ) on the right side of every Data Job/Job Failures assets you aim to modify. Then, select the 'View Monitor' option.

<figure><img src="/files/cVUIkxd3yudRg8XE7S1h" alt=""><figcaption></figcaption></figure>

Click on 'Modify' button to enable the form.

<figure><img src="/files/04G8tAJVAOg8prvy8tgj" alt=""><figcaption><p>Modify monitor modal form</p></figcaption></figure>

Then, you can modify the Get Notified and Incident Levels settings of the monitor. Click 'Save' and you will see a toast message that says, "Monitor updated successfully".

<figure><img src="/files/MZKzhLVh8ZZO7KmEqxEU" alt=""><figcaption></figcaption></figure>

You may also disable or enable the monitor:

* if disabled, the monitor will be Inactive.
* if enabled, the monitor will be Active.

<figure><img src="/files/apB10i0dGsVNp15SYIg2" alt=""><figcaption><p>Disabled monitor</p></figcaption></figure>

<figure><img src="/files/1S8i9zRMQ7QVWv29Mj2c" alt=""><figcaption><p>Enabled monitor</p></figcaption></figure>

### Notifications and Custom Alerts

Locate the "Notify default channels" toggle to receive automatic notifications for detected incidents simply by just switching the toggle to the "on" position.

<figure><img src="/files/Qt65QyTjhFUgtPiu4wqZ" alt=""><figcaption></figcaption></figure>

To personalize your alert settings in the monitor, especially if you wish to be notified of any anomalies or incidents via specific channels like Email or Slack, simply click on the "+ Add other custom alerts" button.

Upon checking the "+ Add other custom alerts" button, an overlay will appear below for you to set up your preferred medias to be notified.

<figure><img src="/files/JNVKRDd5YSCxF3OY8CDe" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Ensure the email you've provided is correct and remember to "Authorize" Slack within the "My Accounts" section under "Config Settings" tab -> "Config Alerts" This will allow you to send custom alerts to your designated channel.
{% endhint %}


# Monitor Behaviour and Design Considerations

How the Data Quality module is designed to behave in specific scenarios, and how to configure monitors to get the most out of each.

The Data Quality module makes a number of deliberate design choices about how monitors train, alert, and respond to configuration changes. Understanding these upfront helps you set up monitors that work reliably from day one.

***

## Sensitivity adapts through use

Rather than asking you to set a sensitivity value at monitor creation — when you have no data to reason from — Decube starts every monitor at the default sensitivity (0) and lets you adjust it based on what you observe in practice.

Once a monitor has been running and producing incidents, you can tune its confidence interval through the incident feedback mechanism. Marking incidents as false positives shifts the model toward fewer alerts; marking missed anomalies shifts it toward more. This means sensitivity improves as your monitors accumulate real operational history.

{% content-ref url="/pages/oHQMBTOmZQevImaIjmyw" %}
[Incident model feedback](/data-quality/data-quality/incident-model-feedback)
{% endcontent-ref %}

***

## The ML model waits for sufficient signal before alerting

To avoid producing unreliable alerts on insufficient data, the ML model requires a minimum of **5 valid data points in the last 30 observations** before it begins raising incidents. On a newly created monitor or a low-volume table, scans may run without producing incidents until this threshold is reached — this is by design.

Once the signal threshold is met, the model activates and begins producing confidence-interval-based alerts.

{% hint style="success" %}
If you need immediate coverage on a low-volume or sparsely-updated table, use On-Demand mode with a manual threshold. Manual thresholds evaluate on every scan without requiring prior observations.
{% endhint %}

{% content-ref url="/pages/12tZSWlWRX4XEXBTD4L6" %}
[How Anomaly Detection Works](/data-quality/anomaly-detection-explained)
{% endcontent-ref %}

***

## Custom SQL gives you full control over thresholds

Custom SQL monitors evaluate threshold conditions you define explicitly, rather than using the ML model. This makes them well-suited to complex business rules and cross-table assertions where the right threshold is deterministic — for example, "zero rows should have a negative price" or "these two tables must always agree."

{% content-ref url="/pages/PWGhO7SkX4GtJu3jfRBX" %}
[Set Up Custom SQL Monitors](/data-quality/how-to-set-up-monitors/custom-sql-monitors)
{% endcontent-ref %}

***

## Retraining gives you a clean, accurate baseline

When you make a significant change to a monitor's configuration — such as changing the scan frequency, row filtering mode, or timestamp column — Decube starts the training process fresh. This ensures the model is trained on data that reflects the monitor's new configuration, rather than producing a baseline that mixes two different measurement approaches.

Because the previous history no longer reflects the new configuration, it is cleared as part of the retrain. Before making any change that triggers a retrain, note or resolve any open incidents you want to preserve.

{% content-ref url="/pages/KgXixMx9V2joZYvvAjL9" %}
[Retraining Monitors](/data-quality/monitor-configuration-settings/retraining-monitors)
{% endcontent-ref %}


# Trusty AI: Overview

Ask Trusty questions on your metadata.

Trusty AI is built as an AI-powered semantic layer for your metadata. It reduces "search friction" by allowing you to identify relevant data assets using natural language instead of manual browsing or strict naming conventions.

{% embed url="<https://youtu.be/Es4esmBBCFg?si=xk_scFtqYttVNKqf>" %}

## Core Capabilities

* **Semantic Discovery:** Find datasets across warehouses using business metrics or plain language.
* **Lineage Analysis:** Map impact pathways and understand data dependencies through simple queries.
* **Operational Health:** Synthesize monitor statuses and incident history to understand data reliability.
* **Data Profiling:** Read column-level statistics — null rates, value ranges, cardinality, and distributions — directly from Decube's profiling data. Compare historical profiles to identify trends or regressions over time, and get recommendations for which monitors to set up based on profiling signals.

To see a list of prompts that you can ask TrustyAI, go to [Suggested Questions to ask Trusty](/trusty-ai/trusty-ai/suggested-questions-to-ask-trusty)

## Skills

Trusty AI can also run **skills** — prebuilt, repeatable expert workflows such as Monitor Recommendation. The same skills are available for download into your own AI client. For the full list and details, see [Skills](/mcp/skills).

## AI Agent Registry

Trusty sits in the **AI Hub** alongside the AI Agent Registry, an inventory of the AI agents running against your data.

{% content-ref url="/pages/Vt6Z6dripaRFTYg56fsv" %}
[AI Agent Registry: Overview](/ai-agent-registry/overview)
{% endcontent-ref %}

## Working with responses

* **Copy anything:** Copy or download a single code block (SQL, Python, JSON) or an entire response in one click.
* **Work with tables:** Copy or download a table as CSV or Markdown, or maximise it into an expanded view for easier reading.
* **Visualise:** Trusty detects when an answer is best expressed visually and renders it inline in the conversation. Use cases such as but not limited to lineage, pipelines, and relationships can be rendered as diagrams that you can copy and download.
* **Regenerate:** Re-run your question for a fresh answer. Trusty adds the new response below the original, so you keep both.

## Accessing Trusty AI

<figure><img src="/files/uW0vACzoqbyJ53wseXoo" alt=""><figcaption></figcaption></figure>

Trusty AI is currently in a BETA phase. If it is enabled for your organization, follow these steps to access it:

1. Log in to your Decube instance.
2. Click **AI Hub** in the left navigation bar. The AI Hub opens on the Trusty chat by default.
3. Type your request or question into the chat interface at the bottom of the screen.

If you do not see AI Hub in your navigation bar, please contact your Account Manager to request access.

## Security and Data Usage

Trusty AI is built with enterprise-grade security as a primary requirement:

* **Metadata Only:** Trusty AI interacts exclusively with extracted metadata. It cannot execute queries against your raw data sources and has no permission to view record-level data.
* **Copying and downloads stay within your permissions:** Copy and download actions only ever contain contents you are already authorised to see under RBAC. Once you download or copy content, it sits outside Decube's controls.
* **Infrastructure:** The assistant utilizes **AWS Bedrock foundation models**.
  * For SaaS customers, this is managed within Decube’s secure AWS environment.
  * For BYOC (AWS) customers, Trusty connects via your own AWS Bedrock API to ensure data stays within your infrastructure.
* **Data Retention:** Decube stores natural language prompts for review and troubleshooting purposes only.

## Usage and Quotas

To ensure system stability, quota limits are applied to every account.

* You can view your remaining quota directly below the search bar in the Trusty AI interface.
* Regenerating a response re-runs your question against the model and counts toward your quota.

## FAQ

1. **Does Trusty AI have access to my actual data records?**

No. Trusty AI only accesses the metadata (table names, column names, descriptions, and profiling statistics) already ingested by Decube. It never sees the individual rows of data in your warehouse.

2. **What happens if Trusty AI provides an incorrect answer?**

As Trusty AI is in a beta phase, it may occasionally generate "hallucinations" or factually incorrect statements. Verify critical information (like lineage dependencies) via the Decube UI. Please report these errors to `support@decube.io`.

3. **Can Trusty AI make changes in Decube?**

Only one: it can create a monitor. Trusty proposes a monitor with a recommended test type and configuration, and creates it only after you approve, and only if your RBAC permissions allow you to create monitors. Every monitor created this way is attributed to you in Decube's audit trail. Everything else is read-only: Trusty cannot update descriptions, change ownership, or modify any other asset or documentation within Decube as of now.

4. **Why can't I access Trusty AI on my BYOC Azure environment?**

Currently, Trusty AI only supports integration via the AWS Bedrock API. Customers on Azure-based self-hosted deployments are not yet eligible for the current phase Trusty AI beta. If you are hosted on Azure and would like to be on the waitlist for this, please express interest to your Account Manager.

5. **Who can see my chat history?**

Chat histories are visible to you to help you revisit previous sessions. However, Decube team may review anonymized logs strictly for troubleshooting and model optimization purposes.


# Suggested Questions to ask Trusty

To get the most out of Trusty AI, use prompts tailored to your specific role and goals.

The following suggestions are categorized by role to help you get started.

{% embed url="<https://youtu.be/4kDq4K3W83E>" %}

#### Data Engineers

Use Trusty AI to perform rapid impact analysis and check the operational status of your pipelines without writing SQL or traversing lineage graphs manually.

* "Show me the upstream sources for the `fct_monthly_sales` table."
* "Are there any active incidents or failed monitors for the Snowflake production warehouse?"
* "List all downstream dashboards that will be affected if I modify the `user_id` column in the staging schema."
* "Identify the owner of the `raw_transactions` table."
* "Has data quality in `dim_products` changed between the last two profiles?"
* "Which columns in `fact_orders` would you recommend adding monitors for, based on the latest profile?"
* "Draw a lineage diagram for the `monthly_sales` table showing its upstream sources."
* "Tabulate the list of failed monitors in the Snowflake production warehouse to CSV."

#### Data Analysts

Use Trusty AI to discover relevant datasets for your reports and understand the quality of the data before you begin your analysis.

* "Find all datasets related to 'Customer Churn' in the marketing database."
* "Can you simplify my `orders` table lineage into logical business domains for easier understanding? Visualise it."
* "What is the null rate for the `email` column in the `leads` table?"
* "What charts are available in the `Campaign Dashboard`?"
* "Compare the row counts of `stg_orders` and `fct_orders` for the last 7 days."
* "What's the typical value range and distribution for `order_amount` in the `fact_orders` table?"
* "Is `customer_id` a reliable column to join on?"
* "Show how `fact_orders` relates to its dimension tables as a diagram."
* "Give me the null rates for the `leads` table so I can download them as CSV."

#### Data Governance Leads

Use Trusty AI to audit your documentation, track PII, and ensure that your data catalog aligns with your business glossary.

* "Which tables in the 'Sales' schema are missing descriptions?"
* "Identify all assets that are currently linked to the 'GDPR' business glossary term."
* "Show me a summary of data profiling results for the `orders` table to check for uniqueness constraints."
* "List all columns in the `customers` table that have been identified as potentially containing sensitive information."
* "Have null rates across the `dim_customers` table increased since the last profile run?"
* "Diagram which tables feed the `Customer 360` dashboard."
* "List all assets linked to the `GDPR` term and tabulate them."

#### Tips for working with Trusty AI

To receive the most accurate information from Trusty AI, follow these guidelines:

* **Use Backticks for Asset Names**: When referencing a specific table, schema, or column, you enclose the name in backticks (e.g., `` `dim_users` ``) to help the model identify it as a technical asset that you're looking for.
* **Be Context-Specific**: Instead of asking "Show me monitors," try "Show me monitors in the Snowflake `Reporting` schema created in the last 30 days."
* **Leverage Follow-up Questions**: Trusty AI maintains session context. If you find a table, you can immediately ask a follow-up like "Now show me the profile statistics for that table."
* **Verify Outputs:** Trusty AI is an assistant designed to accelerate your workflow. For critical production changes, always verify the lineage or monitor status via the All Monitors or Lineage tabs in the Decube UI.


# LLM Security & Privacy

Understand how Trusty handles LLM-based interactions, what data is passed to the model, and the security and privacy controls in place.

Trusty is built on enterprise-grade LLM infrastructure with security and privacy controls at every layer. This page explains how data flows through Trusty's AI features, what the model can and cannot access, and how your data is protected.

{% hint style="info" %}
**Infrastructure note:** In the current phase, Trusty connects via the AWS Bedrock API. Future phases may include support for additional cloud provider APIs, including Azure and GCP.
{% endhint %}

## LLM Hosting

Trusty is powered by the Claude family of models (ranging from Haiku to Opus), accessed through Amazon Bedrock — AWS's fully managed AI service.

Your content is encrypted at rest and in transit, and stays within the AWS region where you are located. Data is never moved outside your region without your knowledge.

Amazon Bedrock maintains the following certifications and compliance standards: FedRAMP Moderate, SOC 1/2/3, ISO 9001/27001/27017/27018/27701, HIPAA, GDPR support, and CSA STAR Level 2.

## LLM Training

Trusty uses pre-trained foundation models from Anthropic (Claude) and other leading AI providers, accessed via Amazon Bedrock.

**Your data is never used to train or fine-tune any model.** Neither AWS nor any third-party model provider — including Anthropic — uses your inputs or outputs to improve their models.

## Data Isolation Across Instances

Each session is fully isolated. Because no user data is used for model training, there is no mechanism by which your data can influence another user's experience. Data processed within Amazon Bedrock is never shared with third-party providers.

## What Data the LLM Can Access

Trusty follows a least-privilege approach to data access during each session:

* **No unrestricted access:** Trusty does not have broad access to your organisation's data or systems.
* **Scoped tooling:** It interacts with data only through narrowly defined tools and interfaces, limited to the task at hand.
* **RBAC enforcement:** Access is governed by Role-Based Access Controls (RBAC). Trusty can only interact with data you are authorised to access — nothing more.
* **No background collection:** There is no ambient data collection occurring outside of active sessions.

## Copy or Download Trusty Output

Trusty lets you copy and download its output such as responses, code blocks, tables, and diagrams.

* **Your responsibility after export:** Once you copy or download content, it leaves Trusty and sits outside Decube's controls. Handle exported files and clipboard content according to your organisation's data-handling policies.

## Summary

| Concern                        | Trusty's Position                                  |
| ------------------------------ | -------------------------------------------------- |
| Where is data hosted?          | AWS region-local, encrypted at rest and in transit |
| Is my data used for training?  | No — never                                         |
| Can my data reach other users? | No — no training means no cross-instance leakage   |
| What data does the LLM see?    | Only what's needed, governed by RBAC               |

## FAQ

#### Does Trusty use my data to train the model?

No. Your inputs and outputs are never used to fine-tune or train any underlying model. This applies to both AWS and Anthropic.

#### Can other users access my data through Trusty?

No. Sessions are fully isolated. Because no user data is used for training, your data cannot be surfaced to another user or instance.

#### What does Trusty actually send to the LLM?

Trusty sends only the metadata that Decube has already ingested — such as table names, column names, descriptions, and profiling statistics. It never accesses raw data records from your warehouse.

#### Which cloud provider powers Trusty's LLM features?

In the current phase, Trusty connects via the AWS Bedrock API. Decube plans to explore support for additional tenant APIs — including Azure and GCP — in future phases. If you are hosted on Azure and would like to be added to the waitlist, reach out to your Account Manager.

#### Where is my data stored during processing?

Data remains within the AWS region where you are located. It is encrypted at rest and in transit, and is never moved outside your region without your knowledge.

#### Who can see my Trusty chat history?

Your chat history is visible to you to help you revisit previous sessions. Decube may review anonymised logs strictly for troubleshooting and quality improvement purposes.

## Further Reading

For detail on Amazon Bedrock's security and compliance posture, refer to the following AWS resources:

* [Amazon Bedrock Security & Compliance](https://aws.amazon.com/bedrock/security-compliance/) — Overview of Bedrock's security features, compliance certifications, and data protection capabilities.
* [Amazon Bedrock Security Documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/security.html) — Technical documentation covering identity and access management, data protection, logging, and more.
* [Amazon Bedrock FAQs – Security](https://aws.amazon.com/bedrock/faqs/) — Answers to common security questions directly from AWS.

For questions or additional compliance documentation (such as AWS compliance certificates), contact <support@decube.io>.


# Fix Incidents with Trusty AI

Hand an incident to Trusty AI in one click to obtain a remediation plan: get the likely root cause and the steps to fix it, without re-describing the problem.

{% hint style="info" icon="telescope" %}
Fixing incidents with Trusty AI is in **limited preview**. We're iterating on it quickly based on how teams actually use it, so tools, capabilities and availability may change at short notice. For access, questions, or feedback on how it's working for you, reach out at <support@decube.io>.
{% endhint %}

When a monitor triggers, Decube tells you what broke. **Fix with Trusty AI** takes the next step: it hands the full incident context to Trusty and asks for the likely root cause and the remediation steps, so you move from alert to fix without piecing the story together yourself.

## Where to find it

The **Fix with Trusty AI** button sits in the top right of the header on the [Incident Details](/data-quality/data-quality/incident-details) page, to the left of **View monitor**.

The button is available on any incident regardless of status.

{% hint style="info" %}
The button only appears for users who have access to Trusty AI. If you can view the incident but do not see the button, your account does not have Trusty AI enabled. Contact us for access.
{% endhint %}

## Before you begin

We highly recommend adding the **Incident Remediation** skill to your Trusty interface.

The skill gives Trusty a fixed procedure for working an incident: which context to pull and in what order, how to judge whether the data actually broke or the monitor is miscalibrated, and how to lay out the fix. Without it, Trusty still answers, but the depth and shape of the answer vary from incident to incident. With it, every incident gets the same structured investigation against the same quality bar.

{% file src="/files/qjKXisVCIh8ufjHcNNdW" %}

For what skills are and how to use them on other surfaces, see [Skills](/mcp/skills).

## How does it work?

1. Select an incident to open the **Incident Details** page.
2. Click **Fix with Trusty AI** in the top right of the incident header.
3. Trusty opens in a new browser tab with the details pre-filled within the prompt box. You can add or remove any context that is not required.
4. Read the remediation plan, then continue the conversation in the same session to dig deeper.

## What gets sent to Trusty

Each click starts a **new** Trusty session. It never posts into a session you already have open, so the incident is answered against a clean context and your existing conversations are left untouched.

The first message is composed for you and carries a structured summary of the incident:

| Field                          | What it contains                                                          |
| ------------------------------ | ------------------------------------------------------------------------- |
| **Incident ID**                | The unique ID of the incident, so Trusty can look up the full record      |
| **Affected asset**             | The table, column, or data job the incident was raised against            |
| **Incident type**              | Freshness, Volume, Field Health, Schema Drift, Custom SQL, or Job Failure |
| **Incident level**             | Info, Warning, or Critical                                                |
| **What triggered it**          | The name and description of the test that breached                        |
| **Observed & expected values** | The value recorded at detection and expected values                       |

Trusty can then pull further detail on its own: the monitor's run history, incident history, lineage, and profiling.

## What Trusty gives back

With the **Incident Remediation** skill in place, Trusty replies in plain language with a full remediation plan:

* **The diagnosis** — whether the data actually broke or the monitor is miscalibrated, and the evidence behind that call.
* **The likely root cause** — what most plausibly caused the breach, including whether this incident is a symptom of an upstream job failure or schema change.
* **The downstream impact** — the dashboards, reports, and jobs consuming the affected data, so you know who to warn.
* **The next steps** — ranked remediation actions with an owner against each, plus advisory SQL you can run yourself to localise the problem.
* **How it gets verified** — what confirms the fix has landed before you close the incident.

The answer is shaped by the incident type. For example:

* A Freshness incident points you at load schedules and upstream job timing
* A Schema Drift incident points you at the structural change and the downstream consumers exposed to it.

You will not get the same generic advice for both.

Where Trusty cannot land on a root cause with reasonable confidence, it says so and tells you what to check next rather than presenting a guess as fact. Nevertheless, as with any Trusty response, verify anything critical.


# Description Autocomplete

The new AI-powered features in the Catalog are designed to help organizations streamline their data curation process by automatically generating consistent, high-quality descriptions and documentation

{% hint style="info" %}
Trusty AI features are opt-in only. For access, please reach out to our Sales team or your Account Manager to enable the Trusty AI trial for your organization account.
{% endhint %}

**Quick video of the Trusty AI’s Autocomplete**

{% embed url="<https://www.loom.com/share/2fdbf8e937784d939a5968285ac8eb2b>" %}

### How to: Write description with Trusty AI

1. Go to Catalog → Asset Details → Schema.
2. Click on `Make changes`

{% hint style="info" %}
Note: The Make Changes button is visible only to users with edit permissions. You must be in the Make Changes state to use Trusty AI.
{% endhint %}

<figure><img src="/files/Vee63GzgkcGPfhlnHt3d" alt=""><figcaption><p>Overview for Make Changes button</p></figcaption></figure>

3. Each field with an empty description will display a magic wand icon.

<figure><img src="/files/cjGpU2W7jdiEXw56lBkw" alt=""><figcaption><p>Magic wand</p></figcaption></figure>

4. Click the `magic wand` icon to auto-generate the description.
   1. A loading state `Trusty AI is writing…` will be shown until the description is ready.
   2. Users can edit the generated description if needed.
5. Click `Request Changes` to submit the description for approval.

### **How to: Re-writing Descriptions with Trusty AI**

1. Go to Catalog → Asset Details.
2. Click on `Make changes`

{% hint style="info" %}
Note: You must be in the Make Changes state to use the Re-write feature.
{% endhint %}

3. Click on a field with an existing description to see the option Re-write with Trusty AI.

<figure><img src="/files/p8724MGjh4VAXjdF2Fmr" alt=""><figcaption><p>How to Re-write with Decube AI</p></figcaption></figure>

4. Click `Re-write with Trusty AI` to generate a revised description.
   1. Review the new description before submitting any changes.
5. If unsatisfied, use the `Regenerate` button to rewrite the description.
   1. This icon only shows up, right after the generation of the autofill description.

<figure><img src="/files/OV7w4pLwnR43rW8XqkRj" alt=""><figcaption><p>Regenerate description Option</p></figcaption></figure>

6. Once satisfied, click `Request Changes` or `Mark Approve` for immediate approval.


# AI Agent Registry: Overview

Track every AI agent running in your organization in one inventory.

{% hint style="info" icon="telescope" %}
The AI Agent Registry is in **limited preview**. We're iterating on it quickly based on how teams actually use it, so tools, capabilities and availability may change at short notice. For access, questions, or feedback on how it's working for you, reach out at <support@decube.io>.
{% endhint %}

Your organization runs AI agents against production data: copilots, retrieval assistants, and scripts that teams built because they needed one. Most were created ad hoc and recorded nowhere central, which makes these critical governance questions hard to answer:

* What agent(s) do I have?
* Who owns these agent(s)?
* What data can it reach? Are among these data PII?

The AI Agent Registry gives you a single inventory of those agents. Each entry carries the governance fields you need before an incident rather than after one: what the agent does, who owns it, which environment it runs in, and whether it is still active.

## Finding the registry

The registry lives in the **AI Hub**, alongside Trusty. Click **AI Hub** in the left navigation bar, then click **Agent Registry** under **AI Governance** in the AI Hub sidebar.

## At a glance

Three cards sit at the top of the registry:

| Card              | What it counts                                                       |
| ----------------- | -------------------------------------------------------------------- |
| **Total Agents**  | Every agent registered in your organization, across all environments |
| **Active Agents** | Agents with an Active status                                         |
| **Owners**        | The number of unique agent owners                                    |

## Filtering the registry

Four tabs sit above the table:

* **All Agents** — every agent registered in your organization.
* **My Agents** — agents you own.
* **Unverified** — agents whose verification is still Unverified.
* **Inactive** — agents with an Inactive status.

Below the tabs, a search input matches on name, description, owner, and source. Five dropdown filters narrow the list further: **Environment**, **Owner**, **Source**, **Status**, and **Verification**.

## What the table shows

| Column            | What it shows                                                                         |
| ----------------- | ------------------------------------------------------------------------------------- |
| **Agent**         | The agent's type badge, its name, and the agent source it belongs to                  |
| **Owner**         | The person accountable for the agent, or **Unassigned** where nobody owns it yet      |
| **Environment**   | Where the agent runs, for example `Production`                                        |
| **Status**        | **Active** or **Inactive**                                                            |
| **Verification**  | **Verified** or **Unverified**                                                        |
| **Last Activity** | Relative time since the agent was last seen, or `-` where there is no activity signal |
| **Actions**       | Per-agent actions, including opening the agent's details                              |

Click any row to open the [agent detail view](/ai-agent-registry/agent-details), which shows the governance fields in full.

## How agents get into the registry

Agents arrive two ways:

* **Automated discovery** — Decube scans a connected source and creates an entry for every agent it finds, so you don't register them by hand.
* **Manual registration** — Manually register agents on platforms Decube can't discover.

Both routes are covered in [Register an Agent](/ai-agent-registry/register-an-agent).

The **Agent sources** button, also at the top right, is where you manage the sources your agents belong to.

Check an agent's **Registration method** on its [detail view](/ai-agent-registry/agent-details) to see which route it took.

## FAQ

#### Who can see the AI Agent Registry?

Anyone in your organization with access to the **AI Hub**.

#### Why don't I see the Agent Registry in the AI Hub sidebar?

During the limited preview, the registry is enabled per organization. If it isn't there, contact us to request access.

#### What does an Unverified agent mean?

Verification is a status the registry tracks alongside Active and Inactive. A newly registered agent starts as Unverified, and you can filter or sort the table on it to find the agents that still need review.

#### Why does an agent show Owner as Unassigned?

Nobody has been recorded as the agent's owner yet. Assign one so the registry can answer "who is accountable for this agent" before you need the answer.

#### Do agents appear in search, lineage, or the catalog?

Not at this stage. Registry entries live in the AI Hub.


# Register an Agent

Get agents into the AI Agent Registry, either by automated discovery via connecting to a source or by registering them manually.

Agents get into the [AI Agent Registry](/ai-agent-registry/overview) two ways: Decube discovers them by scanning a connected source, or you register them manually.

## Automated discovery

An inventory that depends on people remembering to fill in a form is only ever as current as the last person who remembered. Automated discovery removes that step: Decube scans a connected platform and creates a registry entry for every agent it finds.

{% hint style="info" icon="telescope" %}
During the limited preview, automated discovery covers the **Snowflake Cortex** platform, and this may be expanded to other platforms. As of now, agents on any other platform can go into the registry through [manual registration](#manual-registration).
{% endhint %}

### How discovery works

Decube scans your connected source and creates or updates a registry entry for each agent it finds. Each entry arrives with the fields the platform exposes, the agent's name, the source it was found in, and its platform.

Once a source is connected, information about the agent is ingested and synced with your connectors once every hour.

### What a discovered agent looks like on arrival

A scan fills in what the platform reports and leaves the rest empty.

Discovery gives you the inventory. Completing the governance record is a steward's responsibility: assign an owner, and set the environment and lifecycle stage the platform couldn't tell us.

## Manual registration

Manual registration is the path for agents Decube can't discover on its own. Every agent belongs to an **agent source** which is the platform it runs on, so create the source first, then register agents against it.

{% hint style="info" %}
An agent source must exist before you can register an agent. If your organization has no agent source yet, the registration form directs you to create one instead.
{% endhint %}

### Create an agent source

Open the registry and click **Agent sources** at the top right to manage the sources in your organization. Create one source per platform you want to track by hand.

The source you pick shows up next to the agent's name in the registry table, so name it after the platform your team would recognise.

### Register the agent

1. Open the **AI Hub** and click **Agent Registry** under **AI Governance**.
2. Click **Register AI Agent** at the top right.
3. Complete the fields and submit. The agent appears in the registry table immediately.

Write descriptions that tell a steward what the agent does with your data rather than how it was built. "Answers finance questions over the reporting warehouse" is more useful than "LangChain RAG service".

### After registering

The agent's **Owner** shows as **Unassigned** until someone is recorded against it. Assign an owner so the registry can answer who is accountable for the agent before you need the answer.

## FAQ

#### Why are most of the governance fields empty on a discovered agent?

A scan fills in only what the platform exposes about an agent. Everything else stays empty until a steward sets it — use **Edit agent** to complete the record.

#### What's the difference between Source and Platform?

**Source** is the Decube source the agent belongs to, and matches the chip next to the agent's name in the registry table. **Platform** is the agent technology it runs on, such as Snowflake Cortex.

#### Can I register an agent that runs outside a connected platform?

Yes. Manual registration is the path for any agent you want to track, wherever it runs. Create an agent source for its platform first.

#### Can I edit an agent after registering it?

Yes. Open the agent from the registry table and update its fields from the [agent detail view](/ai-agent-registry/agent-details).

#### What happens to the agent source if I register no agents against it?

Nothing. The source stays in your organization as an empty source until you register an agent against it.


# Agent Details

Review an agent's governance fields, the systems it connects to, and how it got into the registry

The agent detail view is where you answer the governance questions the registry table only summarises: who owns this agent, which platform it runs on, what it connects to, and how it got into the registry in the first place.

Open it by clicking an agent row in the registry table, or from the **Actions** menu on that row.

## Viewing Agent Details

The agent's name sits at the top of the page with three badges beside it:

* **Status** — Active or Inactive
* **Verification** — Verified or Unverified
* **Risk level** — the risk rating recorded against the agent, for example `low risk`

Click **Edit agent** at the top right to change the agent's fields. The agent's description sits directly below the header, or **No description provided.** where nobody has written one.

## Governance

The **Governance** section holds the fields a steward needs when reviewing an agent:

| Field                   | What it shows                                                                                                                                 |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| **Owner**               | The person accountable for the agent, or **Unassigned**                                                                                       |
| **Source**              | The agent source the agent belongs to                                                                                                         |
| **Platform**            | The platform the agent runs on, for example Snowflake Cortex                                                                                  |
| **Model**               | The model the agent runs on                                                                                                                   |
| **Environment**         | Where the agent runs, for example `Production`                                                                                                |
| **Lifecycle stage**     | Where the agent sits in its lifecycle                                                                                                         |
| **Agent type**          | The kind of agent, for example Assistant                                                                                                      |
| **Registration method** | How the entry got into the registry — either registered manually, or discovered by a scan, with the source scanned and when it was first seen |

Fields nobody has filled in show a dash (—). A discovered agent arrives with whatever the platform exposes and leaves the rest empty until a steward sets it, so expect gaps on a freshly scanned agent.

## Connected Systems

**Connected Systems** lists what the agent reaches.

Each entry shows the connected asset alongside its type and how the link was established, for example `source · declared`.

Where nothing has been resolved for an agent, the section shows an empty state. Read that as "we don't know what this agent reaches", not "this agent reaches nothing".

## Activity

The **Activity** section shows three values:

| Field              | What it shows                                                 |
| ------------------ | ------------------------------------------------------------- |
| **First seen**     | When the agent first appeared in the registry                 |
| **Last activity**  | The agent's most recent activity signal.                      |
| **Open incidents** | The number of open Decube incidents associated with the agent |

## FAQ

#### Why are most of the governance fields empty on a discovered agent?

A scan fills in only what the platform exposes about an agent. Everything else stays empty until a steward sets it — use **Edit agent** to complete the record.

#### Why does Last activity show a dash?

The platform hasn't given Decube a last-run signal for that agent. It does not mean the agent is idle.

#### What's the difference between Source and Platform?

**Source** is the Decube source the agent belongs to, and matches the chip next to the agent's name in the registry table. **Platform** is the agent technology it runs on, such as Snowflake Cortex.


# MCP: Overview

Understand why Decube built an MCP server, the problems it solves, and how data teams use it today.

{% hint style="info" icon="telescope" %}
The Decube MCP server is in **limited preview**. We're iterating on it quickly based on how teams actually use it, so tools, capabilities and availability may change at short notice. For access, questions, or feedback on how it's working for you, reach out at <support@decube.io>.
{% endhint %}

{% embed url="<https://youtu.be/HIqJWIj4ENU?si=q-kQYtgGDBePEnto>" %}

Your data team already works inside AI tools. Data engineers build and change pipelines from Claude Code and Cursor. Analysts answer business questions from Claude, ChatGPT, or Gemini.

The Decube MCP (Model Context Protocol) server exposes Decube's catalog, lineage, and data quality context as tools that any MCP-compatible AI client can call directly, so your team gets Decube's trust signals at the moment they're writing a query, changing a pipeline, or building a report, not after they open a separate browser tab.

## The problem

AI assistants answer data questions without the quality or lineage context Decube already has. A team ships a report on a deprecated table, or a pipeline change breaks three downstream dashboards, and the AI assistant never knew to warn them because Decube's context wasn't in reach.

This shows up differently depending on your role:

* **Data engineers** keep switching out of their AI tool to check a table's lineage before changing a pipeline, or skip the check entirely and break downstream dependencies they didn't know existed.
* **Data analysts** keep switching out of their AI tool to verify a table's freshness and quality in Decube, or skip the check entirely and build on data they haven't verified. A wrong table surfaces downstream as a broken report or a bad decision.
* **Data governance stewards** at regulated organizations keep catalog metadata complete by hand, asset by asset. Coverage never finishes, which is an audit risk.
* **Heads of Data / CDOs** at regulated organizations can't easily tell whether their oraganisation is ready to withstand scrutiny under frameworks like Indonesia's OJK or Australia's APRA CPS 230.

## Common use cases

### 1. Find trusted data, fast.

Before an analyst builds a report, they need to know which table is the right one and whether it's fresh, complete, and free of open incidents. The MCP server answers this from inside the chat tool an analyst already uses, without a detour through the Decube UI.

> "What tables do we have on customer churn, and can I trust the most recent one?"

### 2. Know the blast radius before you change anything.

A data engineer editing a pipeline in Cursor or Claude Code can check lineage and downstream dependencies before making a change, so they see the blast radius without leaving their editor.

> "What depends on the created\_at column of the sales table before I change it?"

### 3. Spot governance gaps systematically.

A data steward can ask their agent to surface assets with no assigned owner, open incidents, or missing classifications, turning a manual audit into a conversational query.

> "Which assets have no assigned owner? Give me a prioritized list."

### 4. Check readiness for a compliance framework.

A Head of Data or CDO at a regulated organization can ask their agent to check the catalog against a specific regulatory framework's data expectations, such as ownership, classification, and incident coverage, and get a straight answer on how exposed they are before an audit.

> "Are we ready for an OJK audit?"

### 5. Curate the catalog at scale, with approval built in.

A data steward can describe, classify, and assign owners across a whole schema in one instruction, with every change attributed to them and routed through their org's approval workflow where one is configured. Weeks of manual curation collapse into minutes, without losing control over what actually changes.

> "Look at `source X` and suggest me descriptions for all tables in `public` schema. Let me review the suggestions. Once I approve the changes, proceed to write them for me."

## Available tools

The MCP server exposes read, write, and action tools. Decube scopes every result to your existing permissions, and checks every write or action against your access before it runs.

### Catalog

| Tool                       | What it does                                   | Type  |
| -------------------------- | ---------------------------------------------- | ----- |
| `search_assets`            | Search assets by name or type (fuzzy or exact) | Read  |
| `get_asset`                | Get full details for a specific asset          | Read  |
| `count_assets`             | Count assets matching a filter                 | Read  |
| `update_asset_description` | Update an asset's description                  | Write |
| `assign_asset_owner`       | Assign a data or business owner to an asset    | Write |

### Lineage

| Tool          | What it does                                 | Type |
| ------------- | -------------------------------------------- | ---- |
| `get_lineage` | Get upstream/downstream lineage for an asset | Read |

### Data quality

| Tool                  | What it does                                          | Type |
| --------------------- | ----------------------------------------------------- | ---- |
| `get_quality_summary` | Get the quality score and health summary for an asset | Read |

### Monitors

| Tool                     | What it does                               | Type   |
| ------------------------ | ------------------------------------------ | ------ |
| `search_monitors`        | Find monitors by name or asset             | Read   |
| `get_monitor_details`    | Get a monitor's configuration              | Read   |
| `get_monitor_history`    | Get a monitor's pass/fail history          | Read   |
| `get_monitor_run_status` | Check the status of a monitor run          | Read   |
| `create_monitor`         | Create a new monitor                       | Write  |
| `update_monitor`         | Update an existing monitor's configuration | Write  |
| `run_monitor`            | Run a monitor on demand                    | Action |
| `enable_disable_monitor` | Enable or disable a monitor                | Action |

### Incidents

| Tool                              | What it does                                       | Type  |
| --------------------------------- | -------------------------------------------------- | ----- |
| `list_incidents_on_asset`         | List incidents raised on an asset                  | Read  |
| `list_incidents_assigned_to_user` | List incidents assigned to a user                  | Read  |
| `get_incident_details`            | Get full detail on an incident                     | Read  |
| `get_incident_history`            | Get an incident's timeline                         | Read  |
| `update_incident`                 | Update an incident's status, assignee, or comments | Write |

### Profiling

| Tool                    | What it does                        | Type   |
| ----------------------- | ----------------------------------- | ------ |
| `get_profiling_list`    | List profiling runs for an asset    | Read   |
| `get_profiling_result`  | Get the results of a profiling run  | Read   |
| `trigger_profiling_run` | Trigger a profiling run on an asset | Action |

### Governance

| Tool                      | What it does                                | Type  |
| ------------------------- | ------------------------------------------- | ----- |
| `list_custom_attributes`  | List custom attributes defined on assets    | Read  |
| `update_custom_attribute` | Update a custom attribute value on an asset | Write |

### Users

| Tool                 | What it does                              | Type |
| -------------------- | ----------------------------------------- | ---- |
| `find_user_by_email` | Look up a user by email                   | Read |
| `find_user_by_id`    | Look up a user by ID                      | Read |
| `find_users_by_name` | Look up users by name                     | Read |
| `list_all_users`     | List all users in the org                 | Read |
| `whoami`             | Return the identity of the connected user | Read |

### Docs

| Tool            | What it does                          | Type |
| --------------- | ------------------------------------- | ---- |
| `search_docs`   | Search Decube's product documentation | Read |
| `get_docs_page` | Fetch a documentation page            | Read |

## How writes and actions stay governed

Every write and action tool checks your access before it runs. You can only change assets and incidents you're already allowed to edit in Decube, and Decube refuses unauthorized attempts with a clear message.


# Quick Start

Get started using Decube MCP under 3 minutes.

## Connect from Claude

1. In Claude, go to **Customize > Connectors**.
2. Select **Add custom connector**.
3. Open the **Add** dropdown menu on the top-right and select **Add custom connector**.
4. Enter the following details:
   * **Name:** type in `Decube MCP` or another name of your preference
   * **Remote MCP server URL:** type in your region's connector URL e.g. `https://connect.apac.decube.io/mcp`

     | Region | Remote MCP server URL                |
     | ------ | ------------------------------------ |
     | US     | `https://connect.us1.decube.io/mcp`  |
     | EU     | `https://connect.eu1.decube.io/mcp`  |
     | APAC   | `https://connect.apac.decube.io/mcp` |
   * Ignore **advanced settings**
5. Click **Add**.
6. Log in to Decube when prompted, and approve access.

Once connected, try one of the [suggested prompts](/mcp/suggested-prompts) to confirm it's working.

## Connect from Cursor

1. In Cursor, go to **Settings > Tools & MCP**.
2. Click **Add Custom MCP**.
3. Paste the following code snippet in the opened `mcp.json` file. Or add it as a new entry if you have existing MCPs connected.

   ```json
   {
     "mcpServers":
     {
       "decube-mcp": // You can change this to any preferred name
         {
           "url": "https://connect.<region>.decube.io/mcp" // Replace this with your region's Remote MCP server URL
         }
     }
   } 
   ```

   | Region | Remote MCP server URL                |
   | ------ | ------------------------------------ |
   | US     | `https://connect.us1.decube.io/mcp`  |
   | EU     | `https://connect.eu1.decube.io/mcp`  |
   | APAC   | `https://connect.apac.decube.io/mcp` |
4. **Save** the file.
5. If you do not see the new MCP added in your interface, restart Cursor. Otherwise, proceed to the next step.
6. You should see your new MCP in the UI indicating that it requires authentication, click **Connect**. You should be redirected to a web page in approximately 30 seconds.
7. Log in to Decube when prompted, and approve access.

Once connected, try one of the [suggested prompts](/mcp/suggested-prompts) to confirm it's working.

## Connect from ChatGPT (OpenAI)

{% hint style="warning" %}
Custom MCP connectors in ChatGPT require **Developer mode**.
{% endhint %}

1. Turn on Developer mode: in ChatGPT, go to **Settings > Security and login** and toggle on **Developer mode**. If you don't see this option, ask your workspace admin to enable it.
2. Go to **Plugins**.
3. Click on the **+** icon.
4. Enter the following details:
   * **Name:** type in `Decube MCP` or another name of your preference
   * **Connection:** type in your region's connector URL e.g. `https://connect.apac.decube.io/mcp`

     | Region | Remote MCP server URL                |
     | ------ | ------------------------------------ |
     | US     | `https://connect.us1.decube.io/mcp`  |
     | EU     | `https://connect.eu1.decube.io/mcp`  |
     | APAC   | `https://connect.apac.decube.io/mcp` |
   * **Authentication**: OAuth
5. Select **I understand and want to continue**.
6. Select **Create**.
7. Wait for 15-30 seconds for it to connect. Log in to Decube when prompted, and approve access.

Once connected, try one of the [suggested prompts](/mcp/suggested-prompts) to confirm it's working.

## Connect from Microsoft Copilot Studio

{% hint style="info" %}
Reference: [Connect your agent to an existing Model Context Protocol (MCP) server](https://learn.microsoft.com/en-us/microsoft-copilot-studio/mcp-add-existing-server-to-agent).
{% endhint %}

1. Go to the **Tools** page.
2. Select **Add a tool > New tool > Model Context Protocol**.
3. Fill in the required fields:
   * **Server name:** type in `Decube MCP` or another name of your preference
   * **Server description:** e.g. `Decube's catalog, lineage, and data quality context`
   * **Server URL:** type in your region's connector URL

     | Region | Remote MCP server URL                |
     | ------ | ------------------------------------ |
     | US     | `https://connect.us1.decube.io/mcp`  |
     | EU     | `https://connect.eu1.decube.io/mcp`  |
     | APAC   | `https://connect.apac.decube.io/mcp` |
4. Under authentication type, select **OAuth 2.0**, then **Dynamic discovery**.
5. Select **Create**.
6. On the **Add tool** dialog, select **Create a new connection**. Log in to Decube when prompted, and approve access.
7. Select **Add to agent** to finish.

Once connected, try one of the [suggested prompts](/mcp/suggested-prompts) to confirm it's working.

## Next Steps

Now that you're connected, try asking Decube MCP a question relevant to your role:

* [Data engineers](/mcp/suggested-prompts#for-data-engineers) — root-causing failures, dependency checks, freshness breaches
* [Data analysts](/mcp/suggested-prompts#for-data-analysts) — trust checks, ownership lookups, metric definitions
* [Data governance or data stewards](/mcp/suggested-prompts#for-data-governance-or-data-stewards) — audit readiness, ownership gaps, PII classification
* [Data and platform leaders](/mcp/suggested-prompts#for-data-and-platform-leaders) — quality health and compliance summaries

See the full [suggested prompts](/mcp/suggested-prompts) page for more examples, including write and action prompts.


# Suggested Prompts & Use Cases

Example prompts that put Decube's catalog, lineage, and quality context to work in your AI client.

These prompts work today over the MCP server's read, write, and action tools. Ask them in plain English from any connected client, use your own asset, schema, or metric names in place of the examples, and Decube returns results scoped to your existing permissions. Write and action prompts ask for your confirmation before anything changes.

## For data engineers

{% embed url="<https://youtu.be/DsN9vrrFfd4?si=vOQFpUNf-jZC3q1Q>" %}

* "Why did the monitor of the `sales` table from the `Snowflake_Data` source fail recently? Give me the root cause using incident and quality signals."
* "What depends on the `created_at` column of the `sales` table before I change it?"
* "Are there any freshness breaches on my assets right now?"
* "List the open critical incidents on the `orders` table today."
* "Trigger a profiling run on the `customer_id` column of the `customers` table."
* "Run the freshness monitor on the `orders` table now."

## For data analysts

{% embed url="<https://youtu.be/RfKbISoRhg0?si=TJsYoX8GStErx0Af>" %}

* "What tables do we have on customer churn?"
* "What is the official definition and owner for the metric `Net Revenue`?"
* "Is the `customers` table fresh and complete before I build a report on it?"
* "Who owns the `sales` schema, and how do I reach them?"
* "Can I trust the numbers in the `Monthly Revenue` dashboard before I present this?"
* "Are there nulls or outliers in the `customer_id` column?"

## For data governance or data stewards

* "Is my organisation ready for an OJK audit? Generate a report."
* "Recommend me monitors for the Snowflake source. Create the monitors once I've approved them."
* "Which assets have no assigned owner? Give me a prioritized list."
* "Show me lineage evidence for Net Revenue in this report."
* "Which of our assets are already classified as PII?"
* "Look at source X and suggest me descriptions for all tables in public schema. Let me review the suggestions. Once I approve the changes, proceed to write them for me."
* "Assign `jane@company.com` as the owner of the `customers` table."

## For data and platform leaders

* "Is my organisation ready for an OJK audit? Generate a report."
* "What is the quality health of my revenue tables from `<source>`?"
* "Are our PII assets classified?"

{% hint style="info" %}
Some of these questions span multiple assets, so your AI client may make several tool calls to build the full answer. Lineage checks in particular return one hop at a time, so ask a follow-up question to walk further upstream or downstream.
{% endhint %}

{% hint style="info" %}
Write and action prompts ask you to confirm before anything changes, and every change is attributed to you in Decube's audit trail.
{% endhint %}


# Skills

Prebuilt Decube skills — repeatable expert workflows you can run in your own AI client or inside Trusty AI.

{% hint style="info" icon="telescope" %}
Skills are in **limited preview**. We're iterating on it quickly based on how teams actually use it, so tools, capabilities and availability may change at short notice. For access, questions, or feedback on how it's working for you, reach out at <support@decube.io>.
{% endhint %}

A Decube skill is a documented, repeatable workflow. Instead of working out which tools to call and in what order, you run the same expert sequence every time you need it.

Each skill is available on two surfaces:

* **Your own AI client:** download the skill definition and drop it into Claude, Cursor, or another MCP-compatible client.
* **Trusty AI:** the same skills surface inside Decube's in-product assistant.

## Available skills

### Monitor Recommendation

* Reads an asset's profiler statistics and catalog metadata
* Proposes a ranked list of the monitors best suited to it
* Creates the ones you approve

**When to use it**

* After connecting a new source or schema with little monitor coverage
* Whenever you want a starting set of monitors on an asset instead of building from scratch
* To improve the data quality coverage of your assets

{% file src="/files/UUhkTI7Jq0FNXOS1LJNo" %}

### Incident Remediation

* Reads the incident, the monitor that fired, and its run history to judge whether the data broke or the rule is miscalibrated
* Walks lineage upstream for the root cause and downstream for the blast radius
* Returns a remediation plan: the diagnosis, the likely root cause, the consumers affected, ranked next steps with named owners, and how the fix gets verified

**When to use it**

* When you need a remediation plan for an incident
* Alongside [Fix with Trusty AI](/trusty-ai/fix-incident-with-trusty) on the Incident Details page
* When a monitor keeps firing and you need to know whether to fix the data or retune the check

{% file src="/files/qjKXisVCIh8ufjHcNNdW" %}

## Using skills in your own AI client

Download the skill definition and add it to your MCP-compatible client, then run it against Decube's tools. You'll need the Decube MCP server connected first: see [Quick Start](/mcp/quick-start) if you haven't set that up yet.

## Using skills in Trusty AI

The same skills are available inside [Trusty AI](/trusty-ai/trusty-ai), no download or client setup required.


# Overview of Asset Types

This page explains the main asset types in the Decube Catalog and how they relate to each other to provide a quick mental model of what lives in the catalog.

The Catalog top panel shows asset-type filters as clickable pills. For example, click the **Schema** pill to show Schema assets such as schemas and folders.

![Asset type pills - UI screenshot](/files/JckOKoHgQu8tfe0TUEPz)

***

### Short summary (at a glance)

* Schema: container for other assets (folders, schemas). Not a data-bearing object.
* Table: data-bearing object (tables, views, streams).
* Source: where data originates (database, BI tool).
* Chart / Dashboard: visualizations built on tables.
* Column: the smallest data element (column, JSON attribute).
* Data Job / Task: executable units that transform or move data.

***

### Catalog asset definitions

* Chart
  * What: a single visualization (e.g., a Looker or Tableau chart).
  * Notes: Charts can appear on multiple Dashboards. They carry metadata such as description, owner, lineage, and documentation.
* Dashboard
  * What: a collection of charts grouped for a specific use case (e.g., executive or product dashboard).
  * Notes: Includes metadata and lineage like Charts.
* Schema
  * What: a container for other assets (for example, schemas and folders).
  * Notes: Schema assets organize items but do not contain data themselves.
* Data Job
  * What: an executable job that consumes, produces, or transforms data (examples: Airflow DAGs, dbt runs, Fivetran jobs).
  * Subtypes: DataJobRun, DataTaskRun.
* Data Task
  * What: a discrete step inside a data job or pipeline (transform, validate, enrich, move).
* Table
  * What: structured data, often tabular (tables, materialized views, streams, or files in object storage).
  * Common subtypes: Table (physical table), View (materialized or virtual), Virtual Table (logical representation, see [Virtual tables](#virtual-tables) below).
* Source
  * What: the origin system for tables and other assets (databases, data warehouses, BI tools).
  * Common subtypes: Database, Business Intelligence (BI) tool.
* Column
  * What: the smallest logical data element (column, JSON attribute).
  * Examples: Column (physical table column), Virtual Column (from a virtual table).

#### Asset hierarchy (simple)

The following Mermaid diagram shows the top-level relationships between the common asset types.

```mermaid
flowchart LR
  %% Source branches for DB/warehouse, ETL, BI and jobs
  Source["Source (DB / Warehouse / ETL / BI)"] --> DBCollection["Schema"]
  DBCollection --> DBDataset["Table / View"]
  DBDataset --> Column["Column"]

  Source --> ETLCollection["Schema"]
  ETLCollection --> VirtualDataset["Table (Virtual)"]
  VirtualDataset --> VirtualColumn["Column (Virtual)"]

  Source --> DataJob["Data Job"]
  DataJob --> DataTask["Data Task"]

  Source --> Chart["Chart"]
  Source --> Dashboard["Dashboard"]

  %% visual style
  classDef asset fill:#cce6ff,stroke:#222,stroke-width:1px,color:#222;
  class Source,DBCollection,DBDataset,Column,ETLCollection,VirtualDataset,VirtualColumn,DataJob,DataTask,Chart,Dashboard asset;
```

***

### Virtual tables

A virtual table is a logical data model that does not exist as a physical table in a source database. Decube surfaces virtual tables in the catalog when it ingests metadata from transformation tools (such as dbt models) or BI tools (such as Tableau data sources).

Virtual tables appear alongside physical tables in the catalog so that you can document, own, and track lineage for the full data layer — not just the raw storage layer. Each virtual table can carry the same enrichment metadata as a physical table: owner, description, custom attributes, DQ monitors, and lineage edges.

| Asset type     | Where it comes from                                               | Example                                           |
| -------------- | ----------------------------------------------------------------- | ------------------------------------------------- |
| Virtual Table  | dbt, Fivetran, Tableau, and other transformation or BI connectors | A dbt model, a Tableau Published Data Source      |
| Virtual Column | A column projected by a Virtual Table                             | A dbt model column, a calculated field in Tableau |

{% hint style="info" %}
Which connectors produce virtual tables depends on the metadata types each connector supports. Check the **Supported Capabilities** section of your connector's page to confirm whether it collects Virtual Table and Virtual Column metadata.
{% endhint %}

***

### Glossary assets

* Glossary: a central repository of business terms and definitions.
* Term: a single glossary entry (business term) that can be linked to assets for context.
* Category: hierarchical grouping of glossary terms for easier navigation.

![Glossary diagram - UI screenshot](/files/LDfbyScjIxzC6L6s7cKU)


# Assets Catalog

Data dictionary for better collaboration across teams.

{% embed url="<https://youtu.be/YWh2v2hGFzU?si=QEEWlh6jHByBDwZ5>" %}

The Catalog shows all data assets within your active data source. You can filter by asset type — tables, columns, reports, and more — and see any open incidents or assigned ownership for each asset.

<figure><img src="/files/TqMg5OuXsVOAf2KPSR1V" alt=""><figcaption><p>Example of a Data Catalog</p></figcaption></figure>

Selecting an asset from the Catalog opens the Asset Details page, where you can review key statistics and lineage. You can also edit enrichment metadata such as owners and descriptions. Explore related pages below:

{% content-ref url="/pages/afmH9yNlnnj5O4kjDWAi" %}
[Updating Asset Details](/catalog/data-catalog/updating-asset-details)
{% endcontent-ref %}

{% content-ref url="/pages/hNMi9YP1MC3YZhWNlWfD" %}
[Add classifications to fields](/catalog/data-catalog/add-classifications-to-fields)
{% endcontent-ref %}

{% content-ref url="/pages/ZRdnhEirjf6oOvayCLu8" %}
[Automated Lineage](/lineage/automated-lineage)
{% endcontent-ref %}


# Updating Asset Details

How to update a table's overview details and column-level metadata in the Catalog.

{% hint style="info" %}
This page covers updating asset details through the UI, one asset at a time. To update metadata for many assets at once, use [Export/Import](/export-import/export-import-overview). To update asset details programmatically, use the [Assets API](/public-api/overview/index/assets).
{% endhint %}

The Asset Details page is where you maintain the metadata that makes an asset trustworthy and discoverable: description, owners, business owners, custom attributes, linked glossary terms, and column-level classifications. This page walks through updating that information from both the Overview tab and the Schema tab.

## Updating the Overview tab

{% embed url="<https://youtu.be/I5kHhdQLes0>" %}

{% hint style="info" %}
You need Edit access for the asset (granted through the "Asset Details" policy) to make changes. [Learn about source-based policies.](/group-access-policies/source-based-policies)
{% endhint %}

The Overview tab shows the asset's description, owners, business owners, ratings, column classifications, linked data products, and linked glossary terms, along with the full list of columns in the table.

1. Open the asset and select the **Overview** tab to review its current description, owners, business owners, and attributes.

<figure><img src="/files/uMi2dADCKALrO2mCvopC" alt=""><figcaption></figcaption></figure>

2. Scroll down to see the complete list of columns linked to the asset.
3. Click **Make changes** to start editing.
4. Open the **Update asset attributes** dropdown to see the editable sections: **Update owners and description**, **Update custom attributes**, and **Manage linked terms**.

<figure><img src="/files/20782qsWKnd39VbK2Btu" alt=""><figcaption></figcaption></figure>

5. Select a section to open its popup and update the relevant details.

### Updating owners and description

Select **Update owners and description** to open a popup where you set the data owner, business owners, and description, then click **Save preferences**.

<figure><img src="/files/wkgSQaEfSClGI4QEpJ3S" alt=""><figcaption></figcaption></figure>

### Adding a custom attribute

1. Select **Update custom attributes**. A dropdown lists the custom attributes available for this asset — select the ones you want to assign.
2. Enter a value for each selected attribute, then click **Save preferences**. The new attribute value appears immediately in the list.

For the full range of supported attribute types and who can apply them, see [Applying custom attributes to catalog assets](/catalog/data-catalog/apply-custom-attributes).

<figure><img src="/files/KQiL5luVjZjfahYJEkkx" alt=""><figcaption></figcaption></figure>

### Linking glossary terms

1. From **Update asset attributes**, select **Manage linked terms**.
2. Search for the term you want to add and select it to link it directly to the asset.

<figure><img src="/files/DfHBF4jSg6R7bcuWi0Ir" alt=""><figcaption></figcaption></figure>

## Updating column-level details in the Schema tab

The Schema tab lists every column in the table and lets you set classifications, linked terms, descriptions, and custom attributes at the column level.

1. Switch to the **Schema** tab.
2. Click the column you want to edit and add or update its classification, linked terms, description, or custom attributes.

For classification-specific steps, see [Add classifications to fields](/catalog/data-catalog/add-classifications-to-fields).

## Submitting your changes

1. Once your edits are complete, click **Request changes** and review the summary.
2. Select a reviewer, or check **Mark as approved immediately** if you have the necessary permissions, then click **Confirm and request changes**.

<figure><img src="/files/s08J8f9YVM0PCs5U1eq0" alt=""><figcaption></figcaption></figure>

Depending on your access controls, the change either applies immediately or goes to the selected reviewer for approval. For more on how approvals work, see [Initiate a change request](/approval-workflow/summary-of-approval-workflow/initiate-a-change-request).

Keeping the Overview and Schema details current on every asset makes your data catalog easier to trust, search, and govern across your team.


# Applying custom attributes

How to view and apply organization-level custom attributes to catalog assets.

### Overview

Apply custom attributes (created by Org admins) to catalog assets. Applying attributes is done from the Asset Details page using the standard edit/change-request workflows.

What this covers

* Apply, edit and remove custom attribute values on individual catalog assets.
* Attributes are displayed on the Asset Details under a "Custom attributes" section or tab.
* Changes follow the normal change-request process (auto-approve depends on access controls).

<figure><img src="/files/K8Qrlzo69ywuKtMvEM4N" alt=""><figcaption></figcaption></figure>

### Supported attribute types

Custom attributes are available in multiple types so you can capture richer, validated metadata. Supported types:

* Text — free-form string values.
* Integer — whole numbers only; optional min/max constraints.
* User — single or multiple user selection from your organization.
* Enum — admin-defined list of allowed values (each value may have an optional color badge).
* Boolean — Yes / No.

The input widget shown when editing an asset is chosen automatically based on the attribute type. To understand how to add new types or manage existing ones, see [Creating & managing custom attributes](/org-settings/custom-attributes).

### Who can apply attributes

Anyone with the Catalog edit/change-request permission for the asset can add or edit custom attribute values.

For the click path — opening an asset, starting a change request, and setting attribute values — see [Adding a custom attribute](/catalog/data-catalog/updating-asset-details#adding-a-custom-attribute) in Updating Asset Details.

If change approval is required, the update goes for review. If auto-approve is enabled, changes apply immediately.

### Example: Add a "Business Domain" value to a table

1. Open the table asset and click **Make changes**.
2. Under **Update custom attributes**, find "Business Domain" and enter "Retail".
3. Submit the change request. The value appears in the Custom attributes tab once the change is applied.

### Viewing attribute values

* Custom attributes are displayed on the Asset Details page beneath the Summary or in a dedicated Custom attributes tab.
* Values are read-only for anyone without edit rights.

### Use cases

* Business unit tagging: Add a "Business Domain" attribute (Marketing, Finance, Retail) to group tables by ownership and context.
* Sensitivity notes: Short free-text notes such as "PII: limited access" (use classification policies for formal controls).

### Related documentation

* [Create & manage attributes (Org settings)](/org-settings/custom-attributes)
* [Change requests & approvals](/approval-workflow/summary-of-approval-workflow/initiate-a-change-request)


# Add classifications to fields

How to add classifications to columns in the Catalog.

Add classifications to your tables so that others in your team can understand the sensitivity and handling requirements of each column within your data source. This functionality is available in the Schema tab within the Asset Details section.

In the Schema tab, you can see each column within your selected table and edit the row via the kebab menu at the end of each row.

<figure><img src="/files/BkLDrwL2aQ8mooBVhT7U" alt=""><figcaption><p>Example of a populated field catalog.</p></figcaption></figure>

{% hint style="info" %}
To make changes to the Schema tab, you need at least Edit access for the asset, which comes from the "Asset Details" policy. [Learn about source-based policies.](/group-access-policies/source-based-policies)
{% endhint %}

You can also add classifications via the rule-based classification tagging. See [Auto-classify data assets](/governance/auto-classify-data-assets).


# Syncing Metadata from Source

Bring metadata maintained in your data sources directly into the Decube Catalog — reducing duplication and keeping your catalog aligned with the source of truth.

Decube can ingest metadata that already exists in your connected data sources and surface it in the Catalog — so your team doesn't have to maintain the same information in two places.

All syncs are one-way (source → Decube) and read-only within the platform. Synced metadata is refreshed on every metadata ingestion run.

The following metadata types can be synced from supported sources:

* [**Descriptions**](/catalog/syncing-metadata-from-source/sync-from-source) — sync asset descriptions (schema, table, column) from Snowflake or Redshift. Requires configuration at the data source level to enable.
* [**Snowflake Tags**](/catalog/syncing-metadata-from-source/sync-snowflake-tags) — automatically ingest Snowflake tags as System-Managed Attributes in the Catalog. Applies to Snowflake only.
* [**Column Keys & Constraints**](/catalog/syncing-metadata-from-source/sync-snowflake-column-constraints) — sync column-level constraints (NOT NULL, UNIQUE, PRIMARY KEY, FOREIGN KEY) from Snowflake datasets. Applies to Snowflake only.


# Descriptions

Reduces manual export/import work and helps data governance teams maintain a single source of truth for asset descriptions by surfacing the descriptions already maintained in your data warehouse.

The Catalog automatically syncs descriptions from your data source tables, columns, and schemas directly into the Decube Catalog. This one-way sync (source → Decube) keeps your cataloged assets aligned with the source-of-truth descriptions in your warehouse.

When connecting a source, you can choose whether to manage metadata descriptions manually within the platform or sync them from the connected data source. This lets teams decide their preferred source of truth for descriptions, reducing confusion and ensuring consistent documentation practices.

### What this feature does

* Allows you to configure description mode at source connection level.
* Adds a single description field in the Catalog UI either:
  * Synced Description (read-only, from source), or
  * Description (editable, managed in Decube)
  * Applies to Schema, Table, and Column descriptions
  * This is a one-way sync: descriptions ingested from source into Decube. Changes made in Decube are not pushed back to the source.

### Supported sources

* Snowflake
* Redshift

We’ll evaluate other sources based on demand.

### Description Modes

{% hint style="info" %}
Only one description field is shown in the UI at a time, based on the selected mode at source level.
{% endhint %}

When connecting to supported data sources, you can choose how descriptions should appear in the Catalog:

#### 1. User-managed Descriptions

* Editable within Decube.
* Displayed as “Description” field in the catalog for column, tables, schemas.
* Supports inline editing and change requests.
* Used when you want to manage descriptions manually in Decube.

#### 2. Synced Descriptions

* Pulled from the data source during metadata ingestion
* Displayed as “Synced Description” field in the catalog for column, tables, schemas.
* Synced descriptions are read-only in the UI — you cannot edit them manually.
* Reflects the current state in your source system

### How to Configure Description Mode

{% hint style="info" %}
This selection does not affect ingestion — both synced and user-managed descriptions are still stored in the backend. Only the UI and CSV export behavior changes.
{% endhint %}

You can configure the description mode when creating or modifying supported data sources (see [Supported Sources](#supported-sources) above):

* Navigate to My Account -> Integrations to connect or modify a Source form.
* Scroll to the section “Select which description should show in the catalog”
* Choose:
  * User-managed descriptions (selected by default)
  * Synced descriptions
* This configuration determines which description is shown for all asset types (schema, table, column) in the catalog.

<figure><img src="/files/Rbkx95mtQBYkxXiBazHI" alt=""><figcaption></figcaption></figure>

* You can change the description mode at any time from the Modify Source form.
* The selected option is applied immediately after changing the mode.
  * The Catalog UI and Export/Import will switch to show only the newly selected description.
* You do not need to re-test or re-validate the connection to modify the setting.
* Descriptions not shown in the UI are still stored in Decube.

### Behavior Across Features

The table below shows how feature behavior changes based on the description mode selected in the source connection form.

| **Feature**                                                           | **User-managed**                  | **Synced**                               |
| --------------------------------------------------------------------- | --------------------------------- | ---------------------------------------- |
| [Catalog UI](#where-can-you-see-these-changes-in-the-catalog-ui)      | Editable Description shown        | Read-only Synced Description shown       |
| [Change Requests](#where-can-you-see-these-changes-in-the-catalog-ui) | Supported                         | Not Supported                            |
| [CSV Export (for table)](#csv-export-import-changes)                  | only `Description` field exported | only `Source Description` field exported |
| [CSV Import](#csv-export-import-changes)                              | Allowed                           | Not allowed (read-only)                  |
| [API](#api)                                                           | Returns both fields               |                                          |

### Where can you see these changes in the Catalog UI

* For Schema/Table descriptions, navigate to Catalog → Data Sources (Redshift & Snowflake) → Asset details → Asset Attributes modal.
* To update a user-managed description, click **Make changes**.
* Only one description is displayed at a time — either `Description` (user-managed) or `Synced description`.
* Synced descriptions are read-only. You can still click **Make changes** to update other attributes in the modal, such as owners.

<figure><img src="/files/w2atu96uHRttJR1ytD3m" alt=""><figcaption></figcaption></figure>

* For Column descriptions, navigate to Catalog → Data Sources → Asset details → Schema Tab.
* To update a user-managed description, click **Make changes**.
* Only one description column is displayed at a time — either `Description` (user-managed) or `Synced description`.
* Synced descriptions are read-only. You can still click **Make changes** to update other attributes in the Schema tab, such as classifications and custom attributes.

<figure><img src="/files/Nnjs0YJETlG1s8amm3bu" alt=""><figcaption></figcaption></figure>

### CSV Export/Import Changes

* The CSV file includes a description column with the following header:
  * User-managed description: `Description`
  * Synced description: `Source Description`

{% hint style="info" %}
Before importing the CSV file into the platform, make sure to remove the `Source Description` column. Since this field is read-only, attempting to import it will result in an “Unknown attribute” error.
{% endhint %}

* The exported CSV includes the description column matching the source's configured description mode.
* To include both descriptions in the exported CSV, check **Include both description and synced description** — both fields are then included in the exported CSV.
* You can update the `Description` field (user-managed) through export/import even when the UI shows the synced description. The value is stored, and becomes visible in the Catalog if you later switch back to user-managed mode.
* This configuration applies only to supported data sources.

<figure><img src="/files/NhmasR0p0DT1IqgkmZlz" alt=""><figcaption></figcaption></figure>

### API

* The API returns both `Description` and `Source Description` for Schema, Table, and Column assets. For details, see the [Asset API reference](/public-api/overview/index/assets).

### Edge Case

Change requests created for description remain valid even if the source is later updated to show synced description. This is an edge case and will not be blocked.

## Beta feedback: we want to hear from you

We’re continuing to improve how descriptions work in the Catalog, and your input can help shape what’s next.

We’d like to hear from you:

* Would you want the ability to push changes made in Decube back to the source system?
* Should you be able to edit or comment on Synced Descriptions, even if they are read-only?
* Do you find the new selection between user-managed and synced descriptions clear and helpful?

Send your thoughts and ideas to <product@decube.io> — we’re listening!


# Snowflake Tags

Automatically ingest and synchronize Snowflake tags as System-Managed Attributes in the Decube Catalog, ensuring a single source of truth for governance and discovery.

Decube reduces manual overhead by automatically syncing Snowflake tags directly into the Data Catalog as System-Managed Attributes. This integration ensures your catalog reflects the exact business vocabulary, security tags, and departmental classifications defined in your data warehouse without any manual configuration.

By inheriting these tags from the source, Decube provides a high-fidelity global view of your data assets, reducing the risk of human error and saving governance teams hours of administrative work.

### What this feature does

* **Automatic Ingestion:** Snowflake tags are automatically ingested during the regular metadata sync.
* **System-Managed Attributes:** Tags appear as read-only attributes, distinct from user-created custom attributes.
* **Unified Discovery:** You can filter and search assets based on the synced Snowflake tags.
* **Rich Data Types:** Supports various tag value types (text, enum/dropdown) as defined in Snowflake.

### Snowflake Tags Sync Behavior

{% hint style="info" %}
Snowflake tags are synced automatically by default. Currently, there is no option to opt-out of this sync.
{% endhint %}

* **Read-Only:** Values for Snowflake tags are managed entirely at the source (Snowflake). They cannot be edited within the Decube UI.
* **Identification:** In the Catalog, these tags are identified as **"SourceName.Database.Schema Snowflake tags"**.

{% hint style="info" %}
This naming convention is used because Snowflake tags are not unique between Snowflake databases and schemas.
{% endhint %}

### Snowflake Tags in the Catalog

Snowflake tags integrate directly into your catalog workflows, providing immediate context for data discovery and governance.

#### Viewing Snowflake Tags

In the **Asset Details and Schema Tab**, Snowflake tags appear alongside other custom attributes in the unified **Attributes** column. This gives you a complete view of all metadata attached to your columns at a glance.

<figure><img src="/files/2coidgClC8YoV0NkAEPF" alt=""><figcaption></figcaption></figure>

#### Catalog Filtering for Discovery

Leverage Snowflake tags to refine your search and locate the exact data assets you need. Use the **Filter by Asset Attributes** workflow to filter by specific tag keys and values.

<figure><img src="/files/WiP7htRTiuG8QX9ghB5u" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
**Limitations: Snowflake Tag Updates**

Changes to tag definitions in Snowflake may require manual intervention to propagate correctly if they conflict with existing values.

**Scenario:** If a Snowflake tag is modified from accepting "any text" to only allowing specific enum values (e.g., changing a free-text tag to only accept "High", "Medium", "Low"), existing objects in Snowflake might still retain old values that are now invalid (e.g., "Urgent").

In this case, Decube will follow the new strict enum definition from Snowflake. As a result, the old value ("Urgent") will no longer be displayed in the Catalog because it does not match the valid list of allowed values. **To fix this, update the tag value on the Snowflake object to match the new enum definition.**
{% endhint %}

***

### Key Benefits

* **Operational Efficiency:** Removes the need to duplicate tagging work between Snowflake and Decube.
* **Risk Mitigation:** Ensures critical tags like PII and Security classifications are instantly visible in the catalog.
* **Single Source of Truth:** Maintains consistency by treating Snowflake as the master record for these specific metadata tags.


# Column Keys & Constraints

Automatically ingest column keys and constraints from Snowflake datasets into the Decube Catalog, giving your team immediate visibility into data model rules and relationships.

Decube now syncs column-level constraints directly from your Snowflake datasets into the Data Catalog. Constraints such as NOT NULL, UNIQUE, PRIMARY KEY, and FOREIGN KEY are captured during metadata ingestion and surfaced alongside each column — giving your data team a complete picture of the underlying data model without switching between tools.

This is a read-only, one-way sync (Snowflake → Decube). Constraints are managed at the source and reflected automatically in the Catalog on each metadata ingestion run.

### What this feature syncs

Decube ingests the following column-level constraint types from Snowflake:

* **NOT NULL** — columns that require a value on every row
* **UNIQUE** — columns where all values must be distinct
* **PRIMARY KEY** — the column (or combination of columns) that uniquely identifies each row in a table
* **FOREIGN KEY** — columns that reference a primary key in another table, establishing referential integrity between datasets

{% hint style="info" %}
Snowflake enforces NOT NULL constraints but treats UNIQUE, PRIMARY KEY, and FOREIGN KEY as informational (not enforced). Decube syncs the constraint definition as declared in Snowflake, regardless of enforcement status.
{% endhint %}

### Where constraints appear in the Catalog

Column constraints are visible in the **Schema Tab** of any Snowflake asset in the Catalog. Each column row displays its associated constraints alongside existing metadata such as data type, description, and classifications.

This makes it straightforward to:

* Identify join keys and referential relationships between datasets
* Spot columns with mandatory value requirements before running monitors or building pipelines
* Understand data model intent without querying Snowflake directly

### Sync behavior

* Constraints are synced automatically as part of the standard Snowflake metadata ingestion — no additional configuration is required.
* Constraint data is refreshed on every ingestion run, keeping the Catalog aligned with your current Snowflake schema definitions.
* Constraints are read-only in the Decube UI and cannot be edited or overridden within the platform.

### Supported source

* Snowflake

### Limitations

* Multi-column (composite) PRIMARY KEY and UNIQUE constraints are surfaced per-column — each participating column is individually marked with the constraint type.
* CHECK constraints are not currently synced.
* Constraints must be declared in Snowflake's information schema to be visible in Decube. Constraints enforced only at the application layer are not captured.


# Profiler

Generate and review historical profiles for your data assets

Profiler helps you understand a table's structure, quality, and column-level patterns from the asset details page. You can generate a new profile, review historical runs, and compare how the table looked at different points in time.

## What you can do in Profiler

Use Profiler to:

* Generate a new profile for a table asset
* View the results from the latest and past profile runs
* Review table-level and column-level statistics

{% hint style="info" %}
Profiler replaces the previous Field Statistics experience and stores historical results so you can access them later.
{% endhint %}

{% embed url="<https://www.loom.com/share/a701ad9d7bce41cf9a8987d2647ca496>" %}

## Generate a new profile

To generate a profile for a table:

1. Open the table asset in the Catalog.
2. Click the **Profile** tab.
3. Click **Generate a new profile** and select a run option.
4. Wait for the profile to complete, or navigate away and return later.

When the profile starts, a new entry appears in the historical profiles list with an **In Progress** status.

There are three run options:

* **Quick Run** — Runs immediately with system defaults.
* **Run with latest settings** — Repeats the most recent configuration used for that asset. Falls back to system defaults if no previous configuration exists.
* **Configure profile run** — Opens a configuration modal where you can choose which columns to profile, apply a date/time filter, and set a sampling strategy.

{% content-ref url="/pages/0F2brgN7N8hZ0mVglRQt" %}
[Configure a profile run](/catalog/profiler/configure-profile-run)
{% endcontent-ref %}

## How profile generation works

* Profiler supports one active profile job per asset at a time.
* The profiling job continues in the background even if you leave the page. You can return later to check the results.
* Very wide tables can hit profiling limits at around 250 columns. This is a rough guideline because numeric columns generate more metrics than non-numeric columns.
* Certain sources have limitations on profiling large or wide tables, which can lead to failed runs. For source-specific details, see [Profiler source support and limitations](/catalog/profiler/profiler-source-support-and-limitations).

## Review profiling results

Each successful profile includes table-level statistics and column-level statistics.

### Table statistics

At the top of the results page, Decube shows summary statistics for the table in separate metric cards.

Depending on the data source, these cards can include:

* Total row count (derived from table metadata; does not change when date/time filters are applied)
* Sampling row count
* Whether the value is calculated exactly or estimated

If a statistic is not supported for the connected source, Decube does not show that metric card.

<figure><img src="/files/Bapajay0R8PymKQSZdrt" alt=""><figcaption><p>Profile results showing table-level metric cards and column-level statistics for a successful run.</p></figcaption></figure>

### Column statistics

The column statistics table shows one row per column and can include:

* Column name
* Data type
* Null count and null percentage
* Unique count and unique percentage
* Additional metrics based on the column type

Depending on the data type and source, additional metrics can include values such as:

* Minimum and maximum values
* Mean, median, and standard deviation
* Zero count and zero percentage
* Minimum, maximum, and average string length
* Empty string count and empty string percentage
* True and false counts and percentages

## Understand unsupported values

Some profile metrics depend on the connected source, the table type, and the column data types.

If a metric is not available for a column or source, Decube shows it as `Unsupported` or does not display the metric at all.

Common reasons include:

* The source does not support a specific metric
* The column data type is not supported for profiling
* The result is based on an estimated row count instead of an exact count
* The source has engine-specific limits for profiling large or wide tables

## Source-specific behavior

Profiler behavior varies by source. For example, some connectors use exact row counts while others rely on estimates from source metadata. Some metrics, such as median or unique counts, are also unavailable on specific engines.

For source-specific limitations, sampling behavior, and unsupported metric types, see [Profiler source support and limitations](/catalog/profiler/profiler-source-support-and-limitations).

## Permissions

To generate a profile, you need permission to create profiles for the asset.

If your access is removed, you can no longer open the Profile tab even if a profile job is already in progress.

For access setup details, see [Source-based Policies](/group-access-policies/source-based-policies).

## FAQ

### Why do I not see any column results?

If the asset has no columns, if none of its columns are supported for profiling, or if the table is wide enough to hit profiling limits, the profile run can return an error state instead of results.

### Why is a newly added column missing from the profile?

Profiler uses the metadata available in Decube at the time the profile runs.

If you add a new column in the source and run a profile before the next metadata ingestion completes, the new column does not appear in that profile result. The column appears only in profiles generated after the next metadata ingestion updates the asset schema in Decube.

### Why does a historical profile show old columns?

Historical profiles are snapshots of the asset at the time the profile was generated. If the schema changes later, older profile runs can still show columns that no longer exist in the current table.

### Can I delete old profiles?

No. Historical profiles are retained and shown in reverse chronological order, with the newest run first.

### Can I download profile results?

Yes. After a profile completes, a **Download CSV** button is available on the results page. This exports the profiling results as a CSV file you can open locally.

### Can I profile only specific columns?

Yes. Select **Configure profile run** when generating a profile and choose **Selected Columns** in the column selection section. This is useful for wide tables or when you only need statistics on a subset of columns.




---

[Next Page](/llms-full.txt/1)

