diff --git a/content/embeds/rc-rdi-create-rdi-workspace.md b/content/embeds/rc-rdi-create-rdi-workspace.md index adb94a13aa..9084d3ec2f 100644 --- a/content/embeds/rc-rdi-create-rdi-workspace.md +++ b/content/embeds/rc-rdi-create-rdi-workspace.md @@ -1,4 +1,4 @@ -To create a Data Integration workspace for an existing [Pro subscription]({{< relref "/operate/rc/databases/create-database/create-pro-database-new" >}}): +To create a Data Integration workspace for an existing [Pro subscription](/content/operate/rc/databases/create-database/create-pro-database-new.md): 1. From the Redis Cloud console, select **Data Integration** from the left-hand menu. If you don't have any workspaces yet, select **Create workspace** to go to the **Create workspace** page. @@ -8,7 +8,7 @@ To create a Data Integration workspace for an existing [Pro subscription]({{< re {{The new workspace button.}} - You can also go to the **Data Integration** tab from your subscription or database page and select **Create workspace** to go to the **Create workspace** page for your subscription. + You can also go to the **Data Integration** tab from your subscription or database page and select **Create workspace** to go to the **Create workspace** page for your subscription. {{The create workspace button.}} @@ -33,4 +33,4 @@ To create a Data Integration workspace for an existing [Pro subscription]({{< re {{The create workspace button.}} -Your workspace will be created in the background. You can select **Create pipeline** to [create your pipeline]({{}}) while the workspace is provisioning, or you can select **Create pipeline later** to go back to the Redis Cloud console. +Your workspace will be created in the background. You can select **Create pipeline** to [create your pipeline](/content/operate/rc/rdi/define.md) while the workspace is provisioning, or you can select **Create pipeline later** to go back to the Redis Cloud console. diff --git a/content/embeds/rc-rdi-secrets-permissions.md b/content/embeds/rc-rdi-secrets-permissions.md index ff70bd2699..3bcfdc04b4 100644 --- a/content/embeds/rc-rdi-secrets-permissions.md +++ b/content/embeds/rc-rdi-secrets-permissions.md @@ -16,4 +16,4 @@ } ``` -After you store this secret, you can view and copy the [Amazon Resource Name (ARN)](https://docs.aws.amazon.com/secretsmanager/latest/userguide/reference_iam-permissions.html#iam-resources) of your secret on the secret details page. Save the secret ARN to use when you [define your source database]({{}}). \ No newline at end of file +After you store this secret, you can view and copy the [Amazon Resource Name (ARN)](https://docs.aws.amazon.com/secretsmanager/latest/userguide/reference_iam-permissions.html#iam-resources) of your secret on the secret details page. Save the secret ARN to use when you [define your source database](/content/operate/rc/rdi/define.md). \ No newline at end of file diff --git a/content/embeds/rdi-db-reqs.md b/content/embeds/rdi-db-reqs.md index e41e39316b..4a6ea69d0a 100644 --- a/content/embeds/rdi-db-reqs.md +++ b/content/embeds/rdi-db-reqs.md @@ -4,10 +4,10 @@ * If you are deploying RDI for a production environment then secure this database with a password and TLS. * Set the database's - [eviction policy]({{< relref "/operate/rs/databases/memory-performance/eviction-policy" >}}) to `noeviction`. Note that you can't set this using - [`rladmin`]({{< relref "/operate/rs/references/cli-utilities/rladmin" >}}), + [eviction policy](/content/operate/rs/databases/memory-performance/eviction-policy.md) to `noeviction`. Note that you can't set this using + [`rladmin`](/content/operate/rs/references/cli-utilities/rladmin/_index.md), so you must either do it using the admin UI or with the following - [REST API]({{< relref "/operate/rs/references/rest-api" >}}) + [REST API](/content/operate/rs/references/rest-api/_index.md) command: ```bash @@ -17,11 +17,11 @@ -X PUT https://:9443/v1/bdbs/ ``` * Set the database's - [data persistence]({{< relref "/operate/rs/databases/configure/database-persistence" >}}) + [data persistence](/content/operate/rs/databases/configure/database-persistence.md) to AOF - fsync every 1 sec. Note that you can't set this using - [`rladmin`]({{< relref "/operate/rs/references/cli-utilities/rladmin" >}}), + [`rladmin`](/content/operate/rs/references/cli-utilities/rladmin/_index.md), so you must either do it using the admin UI or with the following - [REST API]({{< relref "/operate/rs/references/rest-api" >}}) + [REST API](/content/operate/rs/references/rest-api/_index.md) commands: ```bash @@ -34,7 +34,7 @@ -H "Content-Type: application/json" \ -X PUT https://:9443/v1/bdbs/ ``` - If you don't have permissions to use AOF persistence, please check the [Using RDI without persistence]({{< relref "/integrate/redis-data-integration/faq#can-i-use-rdi-without-persistence-enabled" >}}) section in the FAQ. + If you don't have permissions to use AOF persistence, please check the [Using RDI without persistence](/content/integrate/redis-data-integration/faq.md#can-i-use-rdi-without-persistence-enabled) section in the FAQ. * **Ensure that the RDI database is not clustered.** RDI will not work correctly if the RDI database is clustered (but note that the target database *can* be clustered without diff --git a/content/embeds/rdi-when-to-use.md b/content/embeds/rdi-when-to-use.md index e1490f887f..3430621cf2 100644 --- a/content/embeds/rdi-when-to-use.md +++ b/content/embeds/rdi-when-to-use.md @@ -10,15 +10,15 @@ RDI is a good fit when: - You want a self-managed solution or AWS based solution. - The source data changes frequently in small increments. - The source database has no more than 20K changes per second. -- RDI throughput during [full sync]({{< relref "/integrate/redis-data-integration/data-pipelines#pipeline-lifecycle" >}}) +- RDI throughput during [full sync](/content/integrate/redis-data-integration/data-pipelines/_index.md#pipeline-lifecycle) stays below 60K records per second, assuming an average record size of 1KB and a pipeline without transformations. -- RDI throughput during [CDC]({{< relref "/integrate/redis-data-integration/data-pipelines#pipeline-lifecycle" >}}) +- RDI throughput during [CDC](/content/integrate/redis-data-integration/data-pipelines/_index.md#pipeline-lifecycle) stays below 20K records per second, assuming an average record size of 1KB and a pipeline without transformations. - The total data size is no larger than 200GB, so a full sync completes in under an hour without exceeding the throughput limits above. RDI can ingest larger datasets, but it will take longer than an hour. - You don’t need to perform join operations on the data from several tables - into a [nested Redis JSON object]({{< relref "/integrate/redis-data-integration/data-pipelines/data-denormalization#joining-one-to-many-relationships" >}}). -- RDI supports the [data transformations]({{< relref "/integrate/redis-data-integration/data-pipelines/transform-examples" >}}) you need for your app. + into a [nested Redis JSON object](/content/integrate/redis-data-integration/data-pipelines/data-denormalization.md#joining-one-to-many-relationships). +- RDI supports the [data transformations](/content/integrate/redis-data-integration/data-pipelines/transform-examples/_index.md) you need for your app. - Your data caching needs are too complex or demanding to implement and maintain yourself. - Your database administrator has reviewed RDI's requirements for the source database and confirmed that they are acceptable. diff --git a/content/operate/rc/rdi/_index.md b/content/operate/rc/rdi/_index.md index 408e87e988..b19f098d98 100644 --- a/content/operate/rc/rdi/_index.md +++ b/content/operate/rc/rdi/_index.md @@ -14,9 +14,9 @@ weight: 38 tocEmbedHeaders: true --- -Redis Cloud now supports [Redis Data Integration (RDI)]({{}}), a fast and simple way to bring your data into Redis from other types of primary databases. +Redis Cloud now supports [Redis Data Integration (RDI)](/content/integrate/redis-data-integration/_index.md), a fast and simple way to bring your data into Redis from other types of primary databases. -A relational database usually handles queries much more slowly than a Redis database. If your application uses a relational database and makes many more reads than writes (which is the typical case) then you can improve performance by using Redis as a cache to handle the read queries quickly. Redis Cloud uses [ingest]({{}}) to help you offload all read queries from the application database to Redis automatically. +A relational database usually handles queries much more slowly than a Redis database. If your application uses a relational database and makes many more reads than writes (which is the typical case) then you can improve performance by using Redis as a cache to handle the read queries quickly. Redis Cloud uses [ingest](/content/integrate/redis-data-integration/_index.md) to help you offload all read queries from the application database to Redis automatically. Using a data pipeline lets you have a cache that is always ready for queries. RDI Data pipelines ensure that any changes made to your primary database are captured in your Redis cache within a few seconds, preventing cache misses and stale data within the cache. @@ -35,7 +35,7 @@ apps with a rapidly-growing number of users; the performance of the main databas but it will soon struggle to handle the increasing demand without a cache. Use the information in the sections below to determine whether RDI is a good fit for your architecture. See also the -[decision tree for using RDI]({{}}) +[decision tree for using RDI](/content/integrate/redis-data-integration/when-to-use.md#decision-tree-for-using-rdi) which presents the considerations in a straightforward question-and-answer format. ```decision-tree @@ -53,13 +53,13 @@ Each source first imports its selected data during the *initial sync* phase, the RDI Cloud uses the Flink processor for all pipelines. -For more info on how RDI works, see [RDI Architecture]({{}}). +For more info on how RDI works, see [RDI Architecture](/content/integrate/redis-data-integration/architecture/_index.md). ### Pipeline security -Data pipelines are set up to ensure a high level of data security. Source database credentials and TLS secrets are stored in AWS secret manager and shared using the Kubernetes CSI driver for secrets. See [Share source database credentials]({{}}) to learn how to share your source database credentials and TLS certificates with Redis Cloud. +Data pipelines are set up to ensure a high level of data security. Source database credentials and TLS secrets are stored in AWS secret manager and shared using the Kubernetes CSI driver for secrets. See [Share source database credentials](/content/operate/rc/rdi/setup.md#share-source-database-credentials) to learn how to share your source database credentials and TLS certificates with Redis Cloud. -Configure connectivity separately for each source. A source can use a public endpoint or [AWS PrivateLink](https://aws.amazon.com/privatelink/), subject to the source-specific requirements in [Prerequisites](#prerequisites). See [Set up connectivity]({{}}) to learn how to connect your PrivateLink to the Redis Cloud VPC. +Configure connectivity separately for each source. A source can use a public endpoint or [AWS PrivateLink](https://aws.amazon.com/privatelink/), subject to the source-specific requirements in [Prerequisites](#prerequisites). See [Set up connectivity](/content/operate/rc/rdi/setup.md#set-up-connectivity) to learn how to connect your PrivateLink to the Redis Cloud VPC. RDI encrypts all network connections with TLS. The pipeline will process data from the source database in-memory and write it to the target database using a TLS connection. There are no external connections to your data pipeline except from Redis Cloud management services. @@ -67,7 +67,7 @@ RDI encrypts all network connections with TLS. The pipeline will process data fr Before you can create a data pipeline, you must have: -- A [Redis Cloud Pro database]({{< relref "/operate/rc/databases/create-database/create-pro-database-new" >}}) hosted on Amazon Web Services (AWS). This will be the target database. +- A [Redis Cloud Pro database](/content/operate/rc/databases/create-database/create-pro-database-new.md) hosted on Amazon Web Services (AWS). This will be the target database. - One or more supported source databases that are publicly accessible or hosted on an AWS EC2 instance, AWS RDS, or AWS Aurora: | Database | Versions | AWS RDS Versions | @@ -84,40 +84,39 @@ Before you can create a data pipeline, you must have: | Snowflake | - | - | -{{< note >}} -Please be aware of the following limitations: - -- The target database must be a Redis Cloud Pro database hosted on Amazon Web Services (AWS). Redis Cloud Essentials databases and databases hosted on Google Cloud do not support Data Integration. -- The target database must use [high availability]({{< relref "/operate/rc/databases/configuration/high-availability" >}}). It can use either single-zone or multi-zone high availability. -- The target database can use TLS, but can not use mutual TLS. -- The target database can't use Active-Active topology. -- If your source database is not publicly accessible, or if it is a MongoDB Atlas or Snowflake database, it must be hosted on AWS. -- You must use a [custom encryption key on AWS](https://docs.aws.amazon.com/kms/latest/developerguide/create-keys.html) to create the instance hosting the database. -- Each pipeline has one target database shared by all of its sources. -- If the source database is not publicly accessible, you must be able to set up AWS PrivateLink to connect your source database to your target database. RDI only works with AWS PrivateLink and not VPC Peering or other private connectivity options. -- Mutual TLS is not supported for AWS RDS and AWS Aurora source databases. -{{< /note >}} +> [!NOTE] +> Please be aware of the following limitations: +> +> - The target database must be a Redis Cloud Pro database hosted on Amazon Web Services (AWS). Redis Cloud Essentials databases and databases hosted on Google Cloud do not support Data Integration. +> - The target database must use [high availability](/content/operate/rc/databases/configuration/high-availability.md). It can use either single-zone or multi-zone high availability. +> - The target database can use TLS, but can not use mutual TLS. +> - The target database can't use Active-Active topology. +> - If your source database is not publicly accessible, or if it is a MongoDB Atlas or Snowflake database, it must be hosted on AWS. +> - You must use a [custom encryption key on AWS](https://docs.aws.amazon.com/kms/latest/developerguide/create-keys.html) to create the instance hosting the database. +> - Each pipeline has one target database shared by all of its sources. +> - If the source database is not publicly accessible, you must be able to set up AWS PrivateLink to connect your source database to your target database. RDI only works with AWS PrivateLink and not VPC Peering or other private connectivity options. +> - Mutual TLS is not supported for AWS RDS and AWS Aurora source databases. ## Get started -To get started fast with RDI on Redis Cloud, see the [RDI Cloud quick start]({{}}) to create a data pipeline between a PostgreSQL source database and a Redis Cloud target database. +To get started fast with RDI on Redis Cloud, see the [RDI Cloud quick start](/content/operate/rc/rdi/quick-start.md) to create a data pipeline between a PostgreSQL source database and a Redis Cloud target database. To create a new data pipeline, you need to: -1. [Create a Data Integration workspace]({{}}) for your Pro subscription. -1. [Prepare each source database]({{}}) and any associated credentials. -1. [Define the source connection and data pipeline]({{}}) by selecting which tables to sync. +1. [Create a Data Integration workspace](/content/operate/rc/rdi/create-workspace.md) for your Pro subscription. +1. [Prepare each source database](/content/operate/rc/rdi/setup.md) and any associated credentials. +1. [Define the source connection and data pipeline](/content/operate/rc/rdi/define.md) by selecting which tables to sync. -Once your data pipeline is defined, you can [view and edit]({{}}) it. +Once your data pipeline is defined, you can [view and edit](/content/operate/rc/rdi/view-edit.md) it. -For complete production setups, including SQL Server failover handling, see [Production use cases]({{}}). +For complete production setups, including SQL Server failover handling, see [Production use cases](/content/operate/rc/rdi/use-cases/_index.md). ## Billing and common questions -See the [RDI Cloud FAQ]({{< relref "/operate/rc/rdi/faq" >}}) for billing examples, reset and flush behavior, and working with multiple sources. +See the [RDI Cloud FAQ](/content/operate/rc/rdi/faq.md) for billing examples, reset and flush behavior, and working with multiple sources. ## Maintenance windows RDI Cloud maintenance follows the same subscription-wide maintenance window as your Redis Cloud Pro subscription. During a maintenance window, your data pipeline may experience brief interruptions as Redis applies updates. -To control when maintenance occurs, [set a manual maintenance window]({{< relref "/operate/rc/subscriptions/maintenance/set-maintenance-windows" >}}) for your Redis Cloud Pro subscription. Any maintenance window you configure applies to both your databases and your RDI data pipeline. +To control when maintenance occurs, [set a manual maintenance window](/content/operate/rc/subscriptions/maintenance/set-maintenance-windows.md) for your Redis Cloud Pro subscription. Any maintenance window you configure applies to both your databases and your RDI data pipeline. diff --git a/content/operate/rc/rdi/create-workspace.md b/content/operate/rc/rdi/create-workspace.md index 5185def009..68f8a63699 100644 --- a/content/operate/rc/rdi/create-workspace.md +++ b/content/operate/rc/rdi/create-workspace.md @@ -38,9 +38,8 @@ There, you'll see your workspace and its pipelines. The **Sources** column lists ## Delete workspace -{{< warning >}} -Make sure to [delete your data pipeline]({{}}) before deleting your workspace. -{{< /warning >}} +> [!WARNING] +> Make sure to [delete your data pipeline](/content/operate/rc/rdi/view-edit.md#delete-pipeline) before deleting your workspace. To delete your workspace, select **Workspace actions > Delete workspace** from your workspace. diff --git a/content/operate/rc/rdi/define.md b/content/operate/rc/rdi/define.md index fd8584882f..390321c20e 100644 --- a/content/operate/rc/rdi/define.md +++ b/content/operate/rc/rdi/define.md @@ -13,7 +13,7 @@ hideListLinks: true weight: 4 --- -After you have [prepared each source database]({{}}) and [created a workspace]({{}}), you can create a pipeline. One pipeline can ingest data from several source databases into one Redis target. +After you have [prepared each source database](/content/operate/rc/rdi/setup.md) and [created a workspace](/content/operate/rc/rdi/create-workspace.md), you can create a pipeline. One pipeline can ingest data from several source databases into one Redis target. In the [Redis Cloud console](https://cloud.redis.io/), open your target database's **Data Integration** tab and select **Add pipeline**. You can also open the workspace from the **Data Integration** page or your subscription's **Data Integration** tab. To continue an existing draft, open its actions menu and select **Resume pipeline setup**. @@ -35,7 +35,7 @@ To create a pipeline: {{The target database list in pipeline Settings.}} 1. Select **Hash** or **JSON** as the **Default data structure**. Transformation jobs can override how individual records are written. -1. If needed, configure **Processor properties**. These apply to the whole pipeline, not to an individual source. See the [processor configuration reference]({{< relref "/integrate/redis-data-integration/reference/config-yaml-reference#processors-data-processing-configuration" >}}). +1. If needed, configure **Processor properties**. These apply to the whole pipeline, not to an individual source. See the [processor configuration reference](/content/integrate/redis-data-integration/reference/config-yaml-reference.md#processors-data-processing-configuration). {{The processor advanced properties editor with key and value fields.}} @@ -66,9 +66,9 @@ Complete the sections and select **Test source**. Correct any reported errors be ### Source connectivity -Choose **AWS Private Link** or **Public Endpoint** for the selected source, according to its [connectivity requirements]({{}}). +Choose **AWS Private Link** or **Public Endpoint** for the selected source, according to its [connectivity requirements](/content/operate/rc/rdi/_index.md#prerequisites). -- For **AWS Private Link**, enter the **Private Link service name** from your [endpoint service]({{< relref "/operate/rc/rdi/setup#set-up-connectivity" >}}). Select **Connect to Private Link** and wait for connectivity to complete. If the connection fails, check the service name and its allowed principal. +- For **AWS Private Link**, enter the **Private Link service name** from your [endpoint service](/content/operate/rc/rdi/setup.md#set-up-connectivity). Select **Connect to Private Link** and wait for connectivity to complete. If the connection fails, check the service name and its allowed principal. {{AWS Private Link connectivity with the service name and Connect to Private Link control.}} @@ -80,7 +80,7 @@ Configure connectivity for each source separately. Sources in the same pipeline ### Secrets -Enter the Amazon Resource Name (ARN) of the selected source's [database credentials secret]({{< relref "/operate/rc/rdi/setup#create-database-credentials-secrets" >}}) in **Credentials Secret ARN**. +Enter the Amazon Resource Name (ARN) of the selected source's [database credentials secret](/content/operate/rc/rdi/setup.md#create-database-credentials-secrets) in **Credentials Secret ARN**. {{The Credentials Secret ARN field, transit security options, and Validate control.}} @@ -99,7 +99,7 @@ Under **Transit security**, select the mode required by your source: {{mTLS transit security with certificate, private key, and optional password secret ARN fields.}} -Select **Validate** to check access to the selected source's secrets. Repeat this for each source. The AWS secret contents and permissions are described in [Share source database credentials]({{< relref "/operate/rc/rdi/setup#share-source-database-credentials" >}}). +Select **Validate** to check access to the selected source's secrets. Repeat this for each source. The AWS secret contents and permissions are described in [Share source database credentials](/content/operate/rc/rdi/setup.md#share-source-database-credentials). ### Source configuration {#source-configuration-section} @@ -113,7 +113,7 @@ Enter the selected source's database settings. The fields depend on the database {{Source-specific database, port, and collector properties.}} -Use **Collector properties** for additional source and sink settings. These settings apply to the selected source. See the [collector source properties]({{< relref "/integrate/redis-data-integration/reference/config-yaml-reference#sourcesadvancedsource-advanced-source-settings" >}}) and [collector sink properties]({{< relref "/integrate/redis-data-integration/reference/config-yaml-reference#sourcesadvancedsink-rdi-collector-stream-writer-configuration" >}}). +Use **Collector properties** for additional source and sink settings. These settings apply to the selected source. See the [collector source properties](/content/integrate/redis-data-integration/reference/config-yaml-reference.md#sourcesadvancedsource-advanced-source-settings) and [collector sink properties](/content/integrate/redis-data-integration/reference/config-yaml-reference.md#sourcesadvancedsink-rdi-collector-stream-writer-configuration). {{The Edit advanced properties control under Collector properties.}} @@ -123,9 +123,8 @@ Use **Collector properties** for additional source and sink settings. These sett Select each source in the **Sources** list and choose the data to ingest from that source. -{{< warning >}} -Do not write data directly to keys managed by RDI. Changes from another application can cause transformation failures or data inconsistencies, and RDI can overwrite them. A pipeline reset does not flush the target database. **Flush target database** is a separate action that deletes all target data, including data written outside RDI. See [Data recovery]({{< relref "/operate/rc/rdi/faq#data-recovery" >}}). -{{< /warning >}} +> [!WARNING] +> Do not write data directly to keys managed by RDI. Changes from another application can cause transformation failures or data inconsistencies, and RDI can overwrite them. A pipeline reset does not flush the target database. **Flush target database** is a separate action that deletes all target data, including data written outside RDI. See [Data recovery](/content/operate/rc/rdi/faq.md#data-recovery). 1. Select a schema in **Schemas** to see its tables. 1. Select the tables to ingest in **Tables**. @@ -150,7 +149,7 @@ The available schema, table, and column controls depend on the source type. Each Transformation jobs are optional. Without a matching job, RDI writes records using the pipeline's default data structure. -1. Select **Upload jobs** to upload the [transformation job files]({{< relref "/integrate/redis-data-integration/data-pipelines/transform-examples" >}}) needed for your selected tables. +1. Select **Upload jobs** to upload the [transformation job files](/content/integrate/redis-data-integration/data-pipelines/transform-examples/_index.md) needed for your selected tables. 1. Check the **Source name** assignment for each job. In a multi-source pipeline, `source.server_name` identifies the source the job reads from. Select the source name you chose during setup. 1. Review each job's validation status and correct errors. @@ -158,7 +157,7 @@ Transformation jobs are optional. Without a matching job, RDI writes records usi 1. Select **Continue to review & deploy**. -For the Flink processor, source matchers can use lists or entries prefixed with `regex:` to select several tables. Jobs must not overlap on the same table. See [Transformation examples]({{< relref "/integrate/redis-data-integration/data-pipelines/transform-examples" >}}). +For the Flink processor, source matchers can use lists or entries prefixed with `regex:` to select several tables. Jobs must not overlap on the same table. See [Transformation examples](/content/integrate/redis-data-integration/data-pipelines/transform-examples/_index.md). ## Review and deploy {#review-and-deploy} @@ -168,4 +167,4 @@ Select **Deploy pipeline** to start the pipeline. Each source performs its initi {{The Deploy pipeline button.}} -Open the pipeline's [Dashboard and Metrics tabs]({{}}) to follow progress for each source. +Open the pipeline's [Dashboard and Metrics tabs](/content/operate/rc/rdi/view-edit.md) to follow progress for each source. diff --git a/content/operate/rc/rdi/faq.md b/content/operate/rc/rdi/faq.md index ada3e8513a..55c03028f3 100644 --- a/content/operate/rc/rdi/faq.md +++ b/content/operate/rc/rdi/faq.md @@ -18,7 +18,7 @@ A pipeline reset clears the internal RDI state for all sources, including their A reset does **not** flush the target Redis database. Records already in the target remain until RDI overwrites or deletes them through normal processing. Keys that are no longer produced by the current dataset or transformations can remain in the target after a reset. For example, changing a transformation's key prefix and resetting creates keys with the new prefix without deleting keys with the old prefix. -See [Reset data pipeline]({{< relref "/operate/rc/rdi/view-edit#reset-data-pipeline" >}}) for the steps. +See [Reset data pipeline](/content/operate/rc/rdi/view-edit.md#reset-data-pipeline) for the steps. ### What happens when I flush the target database? @@ -26,13 +26,13 @@ See [Reset data pipeline]({{< relref "/operate/rc/rdi/view-edit#reset-data-pipel Flushing does not clear RDI's saved source positions. If you only start the pipeline afterwards, it resumes from those positions. It does not automatically reload records that have not changed in the source. -See [Flush the target database]({{< relref "/operate/rc/rdi/view-edit#flush-the-target-database" >}}). +See [Flush the target database](/content/operate/rc/rdi/view-edit.md#flush-the-target-database). ### How do I reload data after a flush? {#reload-after-flush} 1. Wait for the flush to finish. -1. [Reset the pipeline]({{< relref "/operate/rc/rdi/view-edit#reset-data-pipeline" >}}) while it is stopped, and wait for the reset to finish. -1. [Start the pipeline]({{< relref "/operate/rc/rdi/view-edit#stop-and-restart-data-pipeline" >}}) to take new snapshots of the selected data from all sources. +1. [Reset the pipeline](/content/operate/rc/rdi/view-edit.md#reset-data-pipeline) while it is stopped, and wait for the reset to finish. +1. [Start the pipeline](/content/operate/rc/rdi/view-edit.md#stop-and-restart-data-pipeline) to take new snapshots of the selected data from all sources. 1. Check each source's initial sync progress and record counts on the **Dashboard** and **Metrics** tabs. Wait for initial sync to finish before relying on the target as a complete copy of the selected data. RDI reloads data available in the source databases using the current dataset and transformation settings. It cannot restore data that existed only in the target. @@ -45,23 +45,23 @@ RDI clears the selected source's internal state and takes a new snapshot of its All records already in the target remain, including records from the reset source. The new snapshot can overwrite that source's records. Resetting a source does not selectively delete its target data. -See [Reset a source]({{< relref "/operate/rc/rdi/view-edit#reset-source" >}}). +See [Reset a source](/content/operate/rc/rdi/view-edit.md#reset-source). ### Can I flush data for a single source? {#flush-one-source} Currently, RDI cannot flush target data for a single source. Connect to the target Redis database and selectively delete the records you want to remove. Identify them from your key naming and transformation rules, and check that other sources do not write to the same keys. Do not use **Flush target database** for this purpose: it deletes data for all sources. -Stop the affected source before deleting its records and allow its pending records to finish processing. Starting the source resumes ingestion, and later changes can recreate records. To reload all its selected data, [reset that source]({{< relref "/operate/rc/rdi/view-edit#reset-source" >}}). +Stop the affected source before deleting its records and allow its pending records to finish processing. Starting the source resumes ingestion, and later changes can recreate records. To reload all its selected data, [reset that source](/content/operate/rc/rdi/view-edit.md#reset-source). ### Does deleting a source delete its target data? -No. Deleting a source removes its pipeline configuration and internal RDI state. Records it already wrote to the target remain. You must remove or reassign transformation jobs that refer to the source before deleting it. See [Remove a source]({{< relref "/operate/rc/rdi/view-edit#remove-source" >}}). +No. Deleting a source removes its pipeline configuration and internal RDI state. Records it already wrote to the target remain. You must remove or reassign transformation jobs that refer to the source before deleting it. See [Remove a source](/content/operate/rc/rdi/view-edit.md#remove-source). ## Upgrades and maintenance ### What happens during an RDI Cloud upgrade? -Redis manages RDI upgrades in Redis Cloud. Maintenance follows your Redis Cloud Pro subscription's [maintenance window]({{< relref "/operate/rc/rdi#maintenance-windows" >}}). +Redis manages RDI upgrades in Redis Cloud. Maintenance follows your Redis Cloud Pro subscription's [maintenance window](/content/operate/rc/rdi/_index.md#maintenance-windows). During an upgrade, monitoring may be temporarily unavailable, and ingestion pauses while the pipeline components restart. A streaming pipeline then resumes from its saved state and processes the changes accumulated during the interruption. @@ -108,4 +108,4 @@ Yes. RDI usage is based on the running collectors and processor replicas, not th Creating a workspace or saving a setup draft does not start RDI usage billing. Billing starts when a pipeline is deployed. Stopping the pipeline reduces usage to the workspace charge. Deleting the deployed pipeline ends new RDI usage when no deployed pipelines remain in the workspace; usage already recorded for the hour can still be billed. -Delete an unused pipeline and then [delete its workspace]({{< relref "/operate/rc/rdi/create-workspace#delete-workspace" >}}) when you no longer need RDI. This does not delete the target Redis database or stop its separate charges. +Delete an unused pipeline and then [delete its workspace](/content/operate/rc/rdi/create-workspace.md#delete-workspace) when you no longer need RDI. This does not delete the target Redis database or stop its separate charges. diff --git a/content/operate/rc/rdi/networking/_index.md b/content/operate/rc/rdi/networking/_index.md index b037c73384..99abe67700 100644 --- a/content/operate/rc/rdi/networking/_index.md +++ b/content/operate/rc/rdi/networking/_index.md @@ -11,6 +11,6 @@ linkTitle: Networking weight: 4 --- -Your Data Integration pipeline runs on Redis Cloud and connects to your source database over [AWS PrivateLink]({{}}). The following guides explain how the network path works and how to keep it available: +Your Data Integration pipeline runs on Redis Cloud and connects to your source database over [AWS PrivateLink](/content/operate/rc/rdi/setup.md#set-up-connectivity). The following guides explain how the network path works and how to keep it available: -- [AWS PrivateLink reference]({{}}): How traffic flows between the pipeline and your database, which address each component sees, and how to keep the connection available when your database fails over. +- [AWS PrivateLink reference](/content/operate/rc/rdi/networking/aws-privatelink.md): How traffic flows between the pipeline and your database, which address each component sees, and how to keep the connection available when your database fails over. diff --git a/content/operate/rc/rdi/networking/aws-privatelink.md b/content/operate/rc/rdi/networking/aws-privatelink.md index d7314ebc55..e7714d135b 100644 --- a/content/operate/rc/rdi/networking/aws-privatelink.md +++ b/content/operate/rc/rdi/networking/aws-privatelink.md @@ -11,7 +11,7 @@ linkTitle: AWS PrivateLink weight: 1 --- -This page explains how a Data Integration pipeline reaches your source database over AWS PrivateLink, and how to keep the connection available when your database fails over. For the steps to create the PrivateLink connection, see [Set up connectivity]({{}}). +This page explains how a Data Integration pipeline reaches your source database over AWS PrivateLink, and how to keep the connection available when your database fails over. For the steps to create the PrivateLink connection, see [Set up connectivity](/content/operate/rc/rdi/setup.md#set-up-connectivity). ## How traffic flows {#how-traffic-flows} @@ -38,7 +38,7 @@ The following table shows which address each component sees: Your database always receives the connection from the NLB's own private IP address, never from a Redis Cloud address. PrivateLink and the NLB rewrite the source address as the traffic passes through (network address translation), so Redis Cloud addresses are never visible anywhere in your network. The only firewall rule your database needs is to allow connections from the NLB's subnets. -Because PrivateLink translates addresses instead of routing between the two networks, the workspace CIDR can overlap with your own VPC or on-premises ranges without any conflict. It only needs to be valid on the Redis Cloud side. See [Create a Data Integration workspace]({{}}) for the workspace CIDR requirements. +Because PrivateLink translates addresses instead of routing between the two networks, the workspace CIDR can overlap with your own VPC or on-premises ranges without any conflict. It only needs to be valid on the Redis Cloud side. See [Create a Data Integration workspace](/content/operate/rc/rdi/create-workspace.md) for the workspace CIDR requirements. ## Connect to a database outside the VPC {#connect-to-a-database-outside-the-vpc} @@ -63,7 +63,7 @@ graph LR To set this up: -- When you [create the network load balancer]({{}}), create a target group with target type **IP addresses** and register the database's IP address. +- When you [create the network load balancer](/content/operate/rc/rdi/setup.md#set-up-connectivity), create a target group with target type **IP addresses** and register the database's IP address. - Make sure the VPC can route to that IP address and port, and that the database's firewall or allow list accepts connections from the NLB's subnets. ## Why failover needs an IP address update {#automate-failover} @@ -74,5 +74,5 @@ To recover, the NLB target group must be updated to point to the new address. Th How you update the target group depends on your database: -- **AWS RDS or Aurora**: Use the Lambda function that responds to RDS failover events and updates the target group automatically. To set it up, see [Set up connectivity]({{}}), select the **AWS RDS or Aurora** tab, and follow **Set up Lambda function connectivity**. +- **AWS RDS or Aurora**: Use the Lambda function that responds to RDS failover events and updates the target group automatically. To set it up, see [Set up connectivity](/content/operate/rc/rdi/setup.md#set-up-connectivity), select the **AWS RDS or Aurora** tab, and follow **Set up Lambda function connectivity**. - **Self-managed database (EC2 or on premises)**: Update the target group from your own failover process. You can trigger the update from your database's failover events or from the NLB's health checks. diff --git a/content/operate/rc/rdi/quick-start.md b/content/operate/rc/rdi/quick-start.md index 159e01ea79..efcafcc91b 100644 --- a/content/operate/rc/rdi/quick-start.md +++ b/content/operate/rc/rdi/quick-start.md @@ -16,17 +16,16 @@ weight: 1 The [`rdi-cloud-automation` GitHub repository](https://github.com/redis/rdi-cloud-automation) contains a Terraform script that quickly sets up a PostgreSQL source database on an EC2 instance and all required permissions and network setup to connect it to a Redis Cloud target database. -{{< note >}} -This guide is for demonstration purposes only. It is not recommended for production use. -{{< /note >}} +> [!NOTE] +> This guide is for demonstration purposes only. It is not recommended for production use. ## Prerequisites To follow this guide, you need to: -1. Create a [Redis Cloud Pro database]({{< relref "/operate/rc/databases/create-database/create-pro-database-new" >}}) hosted on Amazon Web Services (AWS). +1. Create a [Redis Cloud Pro database](/content/operate/rc/databases/create-database/create-pro-database-new.md) hosted on Amazon Web Services (AWS). - Turn on Multi-AZ replication and [manually select the availability zones]({{< relref "/operate/rc/databases/configuration/high-availability#availability-zones" >}}) when creating the database. + Turn on Multi-AZ replication and [manually select the availability zones](/content/operate/rc/databases/configuration/high-availability.md#availability-zones) when creating the database. 1. Install the [AWS CLI](https://aws.amazon.com/cli/) and set up [credentials for the CLI](https://docs.aws.amazon.com/cli/latest/userguide/cli-configure-sso.html). @@ -34,13 +33,13 @@ To follow this guide, you need to: ## Create a data integration workspace -Before you can create your first Data Integration pipeline for a Redis Cloud subscription, you must first deploy the cloud infrastructure needed to host the pipeline and run the workers associated with the pipeline. In Redis Cloud, this is called a **Workspace**. See [Create and manage Data Integration workspace]({{}}) for more information. +Before you can create your first Data Integration pipeline for a Redis Cloud subscription, you must first deploy the cloud infrastructure needed to host the pipeline and run the workers associated with the pipeline. In Redis Cloud, this is called a **Workspace**. See [Create and manage Data Integration workspace](/content/operate/rc/rdi/create-workspace.md) for more information. {{< embed-md "rc-rdi-create-rdi-workspace.md" >}} ## Get required ARNs -This example creates one PostgreSQL source. You can [add more sources]({{< relref "/operate/rc/rdi/view-edit#add-source" >}}) after the pipeline is running. +This example creates one PostgreSQL source. You can [add more sources](/content/operate/rc/rdi/view-edit.md#add-source) after the pipeline is running. 1. On the [Redis Cloud console](https://cloud.redis.io/), open your target database's **Data Integration** tab and select **Add pipeline**. @@ -144,13 +143,12 @@ The following example shows the metrics for one selected source in a pipeline wi {{Metrics for one selected source in a pipeline, including snapshot progress and per-table record counts.}} -See [View and edit data pipeline]({{}}) for source actions, dataset changes, and monitoring. +See [View and edit data pipeline](/content/operate/rc/rdi/view-edit.md) for source actions, dataset changes, and monitoring. ## Delete sample resources -{{< warning >}} -Make sure to [delete your data pipeline]({{}}) before deleting the sample resources. -{{< /warning >}} +> [!WARNING] +> Make sure to [delete your data pipeline](/content/operate/rc/rdi/view-edit.md#delete-pipeline) before deleting the sample resources. To delete the sample resources created by Terraform, run: diff --git a/content/operate/rc/rdi/rds-proxy.md b/content/operate/rc/rdi/rds-proxy.md index 64ef8ca50b..cfa0a29389 100644 --- a/content/operate/rc/rdi/rds-proxy.md +++ b/content/operate/rc/rdi/rds-proxy.md @@ -14,13 +14,12 @@ hideListLinks: true weight: 99 --- -{{}} -We do not recommend using RDS Proxy for RDI connections. The [Lambda function approach]({{< relref "/operate/rc/rdi/setup#setup-lambda-function" >}}) provides better failover handling and is the recommended solution for production environments. - -Additionally, RDS Proxy does not work with RDS PostgreSQL and Aurora PostgreSQL because it does not support PostgreSQL logical replication. - -Only use RDS Proxy if you have specific requirements that necessitate it. -{{}} +> [!WARNING] +> We do not recommend using RDS Proxy for RDI connections. The [Lambda function approach](/content/operate/rc/rdi/setup.md#setup-lambda-function) provides better failover handling and is the recommended solution for production environments. +> +> Additionally, RDS Proxy does not work with RDS PostgreSQL and Aurora PostgreSQL because it does not support PostgreSQL logical replication. +> +> Only use RDS Proxy if you have specific requirements that necessitate it. ## Overview @@ -94,22 +93,21 @@ Replace `` with the endpoint of your RDS Proxy. Save this IP add ## Configure the Network Load Balancer -When you [create the Network Load Balancer]({{< relref "/operate/rc/rdi/setup#create-network-load-balancer-rds" >}}), use the RDS Proxy IP address instead of the database IP address: +When you [create the Network Load Balancer](/content/operate/rc/rdi/setup.md#create-network-load-balancer-rds), use the RDS Proxy IP address instead of the database IP address: 1. In **Register targets**, enter the static IP address of your RDS Proxy (obtained in the previous step). 2. Enter the port number where your RDS Proxy is exposed. 3. Select **Include as pending below**. -4. Complete the remaining Network Load Balancer setup as described in the [main setup guide]({{< relref "/operate/rc/rdi/setup#create-network-load-balancer-rds" >}}). +4. Complete the remaining Network Load Balancer setup as described in the [main setup guide](/content/operate/rc/rdi/setup.md#create-network-load-balancer-rds). ## Next steps After setting up RDS Proxy and the Network Load Balancer: -1. [Create an endpoint service]({{< relref "/operate/rc/rdi/setup#create-endpoint-service-rds" >}}) through AWS PrivateLink. -2. [Share your source database credentials]({{< relref "/operate/rc/rdi/setup#share-source-database-credentials" >}}) with Redis Cloud. -3. Continue with the [RDI pipeline configuration]({{< relref "/operate/rc/rdi/define" >}}). +1. [Create an endpoint service](/content/operate/rc/rdi/setup.md#create-endpoint-service-rds) through AWS PrivateLink. +2. [Share your source database credentials](/content/operate/rc/rdi/setup.md#share-source-database-credentials) with Redis Cloud. +3. Continue with the [RDI pipeline configuration](/content/operate/rc/rdi/define.md). -{{}} -When using RDS Proxy, you do not need to set up the Lambda function for failover handling, as the proxy provides a static endpoint. -{{}} +> [!NOTE] +> When using RDS Proxy, you do not need to set up the Lambda function for failover handling, as the proxy provides a static endpoint. diff --git a/content/operate/rc/rdi/setup.md b/content/operate/rc/rdi/setup.md index 56a5712b0d..cd2fe743f6 100644 --- a/content/operate/rc/rdi/setup.md +++ b/content/operate/rc/rdi/setup.md @@ -15,23 +15,23 @@ weight: 3 ## Prepare source database -Prepare every source database before adding it to a pipeline. Each source needs its own change data capture (CDC) configuration, connectivity, and credentials. See [Prerequisites]({{}}) for a list of supported source databases and database versions. +Prepare every source database before adding it to a pipeline. Each source needs its own change data capture (CDC) configuration, connectivity, and credentials. See [Prerequisites](/content/operate/rc/rdi/_index.md#prerequisites) for a list of supported source databases and database versions. -See [Prepare source databases]({{}}) to find steps for your database type: -- [MongoDB Atlas]({{}}) -- [Snowflake]({{}}) +See [Prepare source databases](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/_index.md) to find steps for your database type: +- [MongoDB Atlas](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/mongodb.md) +- [Snowflake](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/snowflake.md) - [Supabase](#supabase) - Hosted on an AWS EC2 instance: - - [MySQL and mariaDB]({{}}) - - [Oracle]({{}}) - - [SQL Server]({{}}) - - [PostgreSQL]({{}}) + - [MySQL and mariaDB](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/my-sql-mariadb.md) + - [Oracle](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/oracle.md) + - [SQL Server](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/sql-server.md) + - [PostgreSQL](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/postgresql.md) - Hosted on AWS RDS or AWS Aurora: - - [AWS Aurora PostgreSQL and AWS RDS PostgreSQL]({{}}) - - [AWS Aurora MySQL and AWS RDS MySQL]({{}}) - - [AWS RDS SQL Server]({{}}) + - [AWS Aurora PostgreSQL and AWS RDS PostgreSQL](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/aws-aurora-rds/aws-aur-pgsql.md) + - [AWS Aurora MySQL and AWS RDS MySQL](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/aws-aurora-rds/aws-aur-mysql.md) + - [AWS RDS SQL Server](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/aws-aurora-rds/aws-rds-sqlserver.md) -See the [RDI architecture overview]({{< relref "/integrate/redis-data-integration/architecture#overview" >}}) for more information about CDC. +See the [RDI architecture overview](/content/integrate/redis-data-integration/architecture/_index.md#overview) for more information about CDC. ### Supabase {#supabase} @@ -39,18 +39,17 @@ You can use a hosted [Supabase](https://supabase.com/) PostgreSQL database as the source for an RDI pipeline on Redis Cloud. The integration was validated with RDI 1.19.0 and hosted Supabase PostgreSQL 17.6. -First, follow the steps in [Prepare Supabase for RDI]({{< relref "/integrate/redis-data-integration/data-pipelines/prepare-dbs/supabase" >}}) +First, follow the steps in [Prepare Supabase for RDI](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/supabase.md) to create the database role, grant table access, account for Row Level Security, and create a publication. Then complete the Redis Cloud-specific steps below. -{{< warning >}} -Supabase AWS PrivateLink connectivity isn't supported. Supabase shares a -Resource Configuration through AWS Resource Access Manager and requires a -Resource-type VPC endpoint. This differs from the AWS PrivateLink -endpoint-service connectivity supported by RDI on Redis Cloud, so you can skip -[Set up AWS Private Link connectivity](#set-up-connectivity) and use the Supabase -public direct database endpoint instead. -{{< /warning >}} +> [!WARNING] +> Supabase AWS PrivateLink connectivity isn't supported. Supabase shares a +> Resource Configuration through AWS Resource Access Manager and requires a +> Resource-type VPC endpoint. This differs from the AWS PrivateLink +> endpoint-service connectivity supported by RDI on Redis Cloud, so you can skip +> [Set up AWS Private Link connectivity](#set-up-connectivity) and use the Supabase +> public direct database endpoint instead. #### Configure public connectivity @@ -133,7 +132,7 @@ After deployment: Supabase logical replication slots retain write-ahead log (WAL) while the pipeline is stopped. Follow the steps in -[Monitor replication slots]({{< relref "/integrate/redis-data-integration/data-pipelines/prepare-dbs/supabase#7-monitor-replication-slots" >}}) +[Monitor replication slots](/content/integrate/redis-data-integration/data-pipelines/prepare-dbs/supabase.md#7-monitor-replication-slots) to monitor retained WAL and prepare for Supabase PostgreSQL upgrades. ## Get cluster account ID @@ -146,7 +145,7 @@ Before you can set up your source connectivity and secrets, you need the AWS Acc {{The Add pipeline control is available while the workspace is being created.}} 1. For a new pipeline, complete **Settings**, including the target database, and select **Continue**. -1. In **Add sources**, select the source type and enter a unique **Source name**. This name identifies the source in the pipeline configuration and transformation jobs. See [Add sources]({{< relref "/operate/rc/rdi/define#pipeline-setup" >}}) for naming rules. +1. In **Add sources**, select the source type and enter a unique **Source name**. This name identifies the source in the pipeline configuration and transformation jobs. See [Add sources](/content/operate/rc/rdi/define.md#pipeline-setup) for naming rules. 1. Select **Continue** to open **Configure source**. 1. Under **Source connectivity**, copy the **Role ARN** and extract its AWS account ID. @@ -161,11 +160,10 @@ Repeat the preparation for each source. Keep track of which endpoint service and ## Set up AWS Private Link connectivity {#set-up-connectivity} -{{< note >}} -If your source database is accessible via a public endpoint and you want to use public connectivity for your data pipeline, proceed to [Share source database credentials](#share-source-database-credentials). -{{< /note >}} +> [!NOTE] +> If your source database is accessible via a public endpoint and you want to use public connectivity for your data pipeline, proceed to [Share source database credentials](#share-source-database-credentials). -If your source database is not accessible via a public endpoint, you need to set up an endpoint service through AWS PrivateLink to be able to connect to it. For how traffic flows over PrivateLink, how to connect to an on-premises database over AWS Direct Connect, and how to keep the connection available during failover, see the [AWS PrivateLink reference]({{}}). +If your source database is not accessible via a public endpoint, you need to set up an endpoint service through AWS PrivateLink to be able to connect to it. For how traffic flows over PrivateLink, how to connect to an on-premises database over AWS Direct Connect, and how to keep the connection available during failover, see the [AWS PrivateLink reference](/content/operate/rc/rdi/networking/aws-privatelink.md). The following diagrams show the network setup for the different database setups: @@ -253,9 +251,8 @@ To connect to your RDS or Aurora database, we recommend using a Lambda function 1. [Create an endpoint service](#create-endpoint-service-rds) through AWS PrivateLink. 1. [Set up Lambda function connectivity](#setup-lambda-function) to route requests to your database. -{{}} -If you have specific requirements that necessitate using RDS Proxy instead of the recommended Lambda function approach, see the [RDS Proxy setup guide]({{< relref "/operate/rc/rdi/rds-proxy" >}}). Note that RDS Proxy is not recommended and does not work with PostgreSQL. -{{}} +> [!NOTE] +> If you have specific requirements that necessitate using RDS Proxy instead of the recommended Lambda function approach, see the [RDS Proxy setup guide](/content/operate/rc/rdi/rds-proxy.md). Note that RDS Proxy is not recommended and does not work with PostgreSQL. ### Create network load balancer {#create-network-load-balancer-rds} @@ -320,9 +317,8 @@ For more details on AWS PrivateLink, see [Share your services through AWS Privat ### Set up Lambda function connectivity {#setup-lambda-function} -{{}} -Setting up the Lambda function is optional but recommended for production environments. The Lambda function provides automatic failover handling and a more robust connection to your RDS or Aurora database. -{{}} +> [!NOTE] +> Setting up the Lambda function is optional but recommended for production environments. The Lambda function provides automatic failover handling and a more robust connection to your RDS or Aurora database. The Lambda function monitors RDS failover events and automatically updates the NLB Target Group to point to the new primary instance's IP address. This ensures RDI reconnects automatically after a failover. @@ -420,9 +416,8 @@ To set up Private Link for a MongoDB Atlas source database: MongoDB Atlas manages its own endpoint service. The flow is a two-way handshake — you get an endpoint service ID from Atlas, give it to Redis Cloud, and then take the VPC Endpoint ID that Redis Cloud returns back to Atlas to complete the connection. -{{< note >}} -Create the Atlas private endpoint in the same AWS region as your Redis Cloud target database. -{{< /note >}} +> [!NOTE] +> Create the Atlas private endpoint in the same AWS region as your Redis Cloud target database. ### Create a private endpoint in MongoDB Atlas @@ -449,9 +444,8 @@ Create the Atlas private endpoint in the same AWS region as your Redis Cloud tar 1. Choose a connection method, then select **Shell**. 1. Copy the connection string shown. - {{< note >}} -Copy the connection string from the **Private Endpoint** connection method only. The standard connection string does not route traffic through the private endpoint. - {{< /note >}} + > [!NOTE] + > Copy the connection string from the **Private Endpoint** connection method only. The standard connection string does not route traffic through the private endpoint. ### Finish pipeline setup @@ -497,9 +491,8 @@ The required secrets depend on your source database's security configuration. Th | mTLS connection |
  • Credentials secret (username and password for the RDI pipeline user)
  • CA Certificate secret (server certificate)
  • Client certificate secret
  • Client key secret
| | mTLS connection with client key passphrase |
  • Credentials secret (username and password for the RDI pipeline user)
  • CA Certificate secret (server certificate)
  • Client certificate secret
  • Client key secret
  • Client key passphrase secret
| -{{< note >}} -{{< embed-md "rdi-tls-secrets.md" >}} -{{< /note >}} +> [!NOTE] +> {{< embed-md "rdi-tls-secrets.md" >}} Select a tab to learn how to create the required secret. @@ -518,9 +511,8 @@ In the [AWS Management Console](https://console.aws.amazon.com/), use the **Serv - `username`: Database username for the RDI pipeline user - `password`: Database password for the RDI pipeline user - {{< note >}} -Snowflake source databases that use user/key pair authentication should only enter the `username` for the database. - {{< /note >}} + > [!NOTE] + > Snowflake source databases that use user/key pair authentication should only enter the `username` for the database. {{< embed-md "rc-rdi-secrets-encryption-permissions.md" >}} @@ -588,4 +580,4 @@ In the [AWS Management Console](https://console.aws.amazon.com/), use the **Serv ## Next steps -After you have prepared connectivity and credentials for each source, resume your pipeline draft from the workspace and complete **Configure source**. Continue with [Create data pipeline]({{}}). +After you have prepared connectivity and credentials for each source, resume your pipeline draft from the workspace and complete **Configure source**. Continue with [Create data pipeline](/content/operate/rc/rdi/define.md). diff --git a/content/operate/rc/rdi/use-cases/_index.md b/content/operate/rc/rdi/use-cases/_index.md index a3946a53fd..cfec9dfa39 100644 --- a/content/operate/rc/rdi/use-cases/_index.md +++ b/content/operate/rc/rdi/use-cases/_index.md @@ -15,4 +15,4 @@ This section describes complete Data Integration setups that customers run in pr Available guides: -- [Connect a data pipeline to on-premises SQL Server and handle failover]({{}}): Connect a pipeline to a self-managed SQL Server Always On availability group and keep it running through failover and disaster recovery events. +- [Connect a data pipeline to on-premises SQL Server and handle failover](/content/operate/rc/rdi/use-cases/sql-server-failover.md): Connect a pipeline to a self-managed SQL Server Always On availability group and keep it running through failover and disaster recovery events. diff --git a/content/operate/rc/rdi/use-cases/sql-server-failover.md b/content/operate/rc/rdi/use-cases/sql-server-failover.md index e68b647d7d..59cfdca830 100644 --- a/content/operate/rc/rdi/use-cases/sql-server-failover.md +++ b/content/operate/rc/rdi/use-cases/sql-server-failover.md @@ -18,7 +18,7 @@ This guide shows how to connect a data pipeline to a self-managed Microsoft SQL The setup in this guide uses: - A SQL Server Always On availability group with a primary replica and one or more readable secondary replicas. The replicas can be in more than one site for disaster recovery. -- A [Network Load Balancer and PrivateLink endpoint service]({{}}) in your AWS VPC. The NLB target group points to the replica that the pipeline should read from, which is usually a readable secondary. If the replicas are on premises, the NLB reaches them over AWS Direct Connect. +- A [Network Load Balancer and PrivateLink endpoint service](/content/operate/rc/rdi/setup.md#set-up-connectivity) in your AWS VPC. The NLB target group points to the replica that the pipeline should read from, which is usually a readable secondary. If the replicas are on premises, the NLB reaches them over AWS Direct Connect. - A Lambda function that updates the NLB target group when the deployment fails over. See [Automate the NLB update](#automate-the-nlb-update). ```mermaid {width="100%"} @@ -52,7 +52,7 @@ The listener can still be useful inside your own network. For example, the Lambd ## Connect to a readable secondary -To make the pipeline read from a readable secondary instead of the primary, add the following property to the collector in your [pipeline configuration]({{}}): +To make the pipeline read from a readable secondary instead of the primary, add the following property to the collector in your [pipeline configuration](/content/operate/rc/rdi/view-edit.md#edit-collector-properties): ```yaml driver.applicationIntent: ReadOnly @@ -104,9 +104,8 @@ When the replica that the pipeline connects to fails, something must point the N - **Event-driven**: When the deployment fails over, a script on any SQL Server node publishes to an SNS topic, and the SNS topic invokes the Lambda function. The Lambda function resolves a known address, such as the availability group listener DNS name, to find the replica to register in the target group. This avoids the detection delay of health checks. - {{< note >}} -Verify what the listener DNS name resolves to in your deployment before you rely on it. A listener DNS name usually resolves to the listener's virtual IP address, which routes to the primary replica. In that case the pipeline reads from the primary after the failover, which works, but the read load moves off the secondary. - {{< /note >}} + > [!NOTE] + > Verify what the listener DNS name resolves to in your deployment before you rely on it. A listener DNS name usually resolves to the listener's virtual IP address, which routes to the primary replica. In that case the pipeline reads from the primary after the failover, which works, but the read load moves off the secondary. - **Health-check driven**: The NLB's own health checks detect the failure and trigger the Lambda function through a CloudWatch alarm. The Lambda function finds the replacement server, for example by an EC2 tag. This needs no changes on the SQL Server side, but detection takes several minutes. @@ -116,7 +115,7 @@ With either trigger, add an email subscription to the SNS topic so that a person Disaster recovery with an availability group is safe for the pipeline. When a secondary is promoted, or you point the NLB at another in-sync replica, the pipeline reconnects and resumes from its saved position with no reset and no data loss. The availability group is responsible for keeping its replicas synchronized, so a promoted replica holds the committed transactions the pipeline has already read. For more information, see [Availability modes for an availability group](https://learn.microsoft.com/en-us/sql/database-engine/availability-groups/windows/availability-modes-always-on-availability-groups) in the SQL Server documentation. -If instead you restore the database from an old backup, rather than recovering it through the availability group, you risk a different situation: the restored database can be behind the position the pipeline has already processed, so the pipeline has nowhere valid to resume from. If you restore from a backup, [reset the pipeline]({{}}) afterward to take a fresh snapshot. +If instead you restore the database from an old backup, rather than recovering it through the availability group, you risk a different situation: the restored database can be behind the position the pipeline has already processed, so the pipeline has nowhere valid to resume from. If you restore from a backup, [reset the pipeline](/content/operate/rc/rdi/view-edit.md#reset-data-pipeline) afterward to take a fresh snapshot. If the restored server is standalone and not yet part of the availability group, enable snapshot isolation before you reset, or the snapshot fails and retries until you do: diff --git a/content/operate/rc/rdi/view-edit.md b/content/operate/rc/rdi/view-edit.md index e2acadbeed..17733ece8a 100644 --- a/content/operate/rc/rdi/view-edit.md +++ b/content/operate/rc/rdi/view-edit.md @@ -44,7 +44,7 @@ A pipeline can be streaming, stopped, or in an error or transitional state. If t 1. On **Dashboard**, select **Add source**. 1. In **Add sources**, select **Add source** in the **Sources** list, choose the new database type, and enter a unique source name. Existing sources remain listed with their connection details read-only. -1. Complete **Configure source**, **Select data**, and **Add transformations** for the new source. See [Create data pipeline]({{< relref "/operate/rc/rdi/define" >}}) for the configuration steps. +1. Complete **Configure source**, **Select data**, and **Add transformations** for the new source. See [Create data pipeline](/content/operate/rc/rdi/define.md) for the configuration steps. 1. Review the changes in **Review & deploy** and select **Deploy pipeline**. {{Add-source wizard showing existing sources and the Add source control.}} @@ -58,7 +58,7 @@ When extending an older single-source pipeline, preserve the original source nam 1. On **Dashboard**, open the source card's **More actions** menu. 1. Select **Stop source**, then confirm with **Stop source**. -Stopping one source pauses data capture from that source without stopping the other sources. Starting it again resumes from its saved position. To resume a source, open its **More actions** menu and select **Start source**, then confirm with **Start source**. A source processes data only while the pipeline is running. See [How does stopping sources affect billing?]({{< relref "/operate/rc/rdi/faq#stopping-and-billing" >}}). +Stopping one source pauses data capture from that source without stopping the other sources. Starting it again resumes from its saved position. To resume a source, open its **More actions** menu and select **Start source**, then confirm with **Start source**. A source processes data only while the pipeline is running. See [How does stopping sources affect billing?](/content/operate/rc/rdi/faq.md#stopping-and-billing). ### Reset a source {#reset-source} @@ -70,7 +70,7 @@ Resetting one source starts a new snapshot and reprocesses its selected data. {{Reset source confirmation explaining that other sources retain their data and the pipeline temporarily stops.}} -The reset clears that source's internal RDI streams, offsets, schema history, rejected records, and processing counters. The whole pipeline and all its sources restart during the reset. Other sources keep their saved positions and resume streaming. All records already in the shared target database remain, including those from the reset source. The new snapshot can overwrite records for that source. See [What happens when I reset one source?]({{< relref "/operate/rc/rdi/faq#reset-one-source" >}}). +The reset clears that source's internal RDI streams, offsets, schema history, rejected records, and processing counters. The whole pipeline and all its sources restart during the reset. Other sources keep their saved positions and resume streaming. All records already in the shared target database remain, including those from the reset source. The new snapshot can overwrite records for that source. See [What happens when I reset one source?](/content/operate/rc/rdi/faq.md#reset-one-source). ### Remove a source {#remove-source} @@ -84,7 +84,7 @@ Before deleting a source, remove or reassign transformation jobs that refer to i Deleting a source removes its data selection and internal RDI state, including streams, offsets, schema history, rejected records, and processing counters. Records already written to the target Redis database remain there. Other sources retain their data. The whole pipeline stops while RDI cleans up the removed source and starts again afterwards. -Source deletion cannot be undone. To remove its records from the target, see [Can I flush data for a single source?]({{< relref "/operate/rc/rdi/faq#flush-one-source" >}}). +Source deletion cannot be undone. To remove its records from the target, see [Can I flush data for a single source?](/content/operate/rc/rdi/faq.md#flush-one-source). ### Change target database @@ -117,13 +117,12 @@ Select a source in the **Sources** list to see its connection status, snapshot p On **Metrics**, select **Connect to Prometheus** to view the available collector and processor endpoints. Add the required endpoints to your Prometheus configuration. Select the relevant source when inspecting per-source metrics. -Prometheus endpoints are exposed on Redis Cloud's internal network. To access this network, enable [VPC peering]({{< relref "/operate/rc/security/vpc-peering" >}}) or [AWS Transit Gateway]({{< relref "/operate/rc/security/aws-transit-gateway" >}}). See [Prometheus and Grafana with Redis Cloud]({{< relref "/integrate/prometheus-with-redis-cloud/" >}}) for more information. +Prometheus endpoints are exposed on Redis Cloud's internal network. To access this network, enable [VPC peering](/content/operate/rc/security/vpc-peering.md) or [AWS Transit Gateway](/content/operate/rc/security/aws-transit-gateway.md). See [Prometheus and Grafana with Redis Cloud](/content/integrate/prometheus-with-redis-cloud/_index.md) for more information. -{{< note >}} -VPC peering and AWS Transit Gateway are the tested and supported methods for accessing Prometheus endpoints. AWS PrivateLink support for this feature is under evaluation and is not currently supported. -{{< /note >}} +> [!NOTE] +> VPC peering and AWS Transit Gateway are the tested and supported methods for accessing Prometheus endpoints. AWS PrivateLink support for this feature is under evaluation and is not currently supported. -For more information about available RDI metrics, see [Observability]({{< relref "/integrate/redis-data-integration/observability" >}}). +For more information about available RDI metrics, see [Observability](/content/integrate/redis-data-integration/observability.md). ## Rejected records @@ -136,14 +135,14 @@ The view shows: - The number of affected tables. - The affected tables and their rejected counts. - Rejected record IDs and rejection times. -- Safe troubleshooting metadata, such as the rejection reason, operation, affected table, and transformation job details when available. See [Using the operation code]({{< relref "/integrate/redis-data-integration/data-pipelines/transform-examples/redis-opcode-example" >}}) for the operation labels. +- Safe troubleshooting metadata, such as the rejection reason, operation, affected table, and transformation job details when available. See [Using the operation code](/content/integrate/redis-data-integration/data-pipelines/transform-examples/redis-opcode-example.md) for the operation labels. Redis Cloud uses the RDI DLQ API to show a sanitized view of rejected records. It does not show the original source record payload or every field stored in the DLQ stream. To inspect the full DLQ entry, connect to the RDI database and read the corresponding DLQ stream directly. -For more information about why records are rejected and how RDI stores them, see [Rejected records]({{< relref "/integrate/redis-data-integration/data-pipelines/rejected-records" >}}). +For more information about why records are rejected and how RDI stores them, see [Rejected records](/content/integrate/redis-data-integration/data-pipelines/rejected-records.md). ## Settings @@ -152,7 +151,7 @@ The **Settings** tab contains the default data structure (**Hash** or **JSON**) {{Pipeline-wide Settings tab with default data structure and processor properties.}} -Select **Edit** to change these settings, then **Save changes** and **Apply and restart**. RDI Cloud uses the Flink processor. Review [processor properties]({{< relref "/integrate/redis-data-integration/reference/config-yaml-reference#processors-data-processing-configuration" >}}) before changing them. +Select **Edit** to change these settings, then **Save changes** and **Apply and restart**. RDI Cloud uses the Flink processor. Review [processor properties](/content/integrate/redis-data-integration/reference/config-yaml-reference.md#processors-data-processing-configuration) before changing them. {{The processor advanced properties editor with key and value fields.}} @@ -164,7 +163,7 @@ Select a source in the **Sources** list to view its connectivity, secret referen 1. Select the source in **Configuration**. 1. Select **Edit collector properties**. -1. Update the [collector source properties]({{< relref "/integrate/redis-data-integration/reference/config-yaml-reference#sourcesadvancedsource-advanced-source-settings" >}}) or [collector sink properties]({{< relref "/integrate/redis-data-integration/reference/config-yaml-reference#sourcesadvancedsink-rdi-collector-stream-writer-configuration" >}}) for that source. +1. Update the [collector source properties](/content/integrate/redis-data-integration/reference/config-yaml-reference.md#sourcesadvancedsource-advanced-source-settings) or [collector sink properties](/content/integrate/redis-data-integration/reference/config-yaml-reference.md#sourcesadvancedsink-rdi-collector-stream-writer-configuration) for that source. {{The advanced properties dialog with separate collector source and sink properties.}} @@ -206,7 +205,7 @@ The **Transformations** tab lists the pipeline's jobs, their source assignments, ### Edit transformations 1. Select **Edit** on **Transformations**. -1. Add, upload, or edit the [transformation jobs]({{< relref "/integrate/redis-data-integration/data-pipelines/transform-examples" >}}) needed for your tables. +1. Add, upload, or edit the [transformation jobs](/content/integrate/redis-data-integration/data-pipelines/transform-examples/_index.md) needed for your tables. 1. Select the **Source name** for each job. 1. Resolve job validation errors and select **Save changes**. 1. Review the restart warning and select **Apply and restart**. @@ -215,7 +214,7 @@ The Flink processor accepts lists and `regex:` entries in source selection patte ## Reset data pipeline -Resetting the whole pipeline clears its internal RDI state for all sources. A running pipeline restarts and takes a new snapshot for every source. A stopped pipeline remains stopped until you [start it](#stop-and-restart-data-pipeline). RDI reprocesses the selected data using the current transformations. Reset alone does not delete records from the target database. See [What happens when I reset the pipeline?]({{< relref "/operate/rc/rdi/faq#reset-pipeline" >}}). +Resetting the whole pipeline clears its internal RDI state for all sources. A running pipeline restarts and takes a new snapshot for every source. A stopped pipeline remains stopped until you [start it](#stop-and-restart-data-pipeline). RDI reprocesses the selected data using the current transformations. Reset alone does not delete records from the target database. See [What happens when I reset the pipeline?](/content/operate/rc/rdi/faq.md#reset-pipeline). 1. Open **Pipeline actions** and select **Reset pipeline**. 1. Review the confirmation and select **Reset data pipeline**. @@ -227,7 +226,7 @@ To re-run the snapshot for just one source while preserving other sources' inter 1. Open **Pipeline actions** and select **Stop pipeline**. 1. Confirm with **Stop pipeline**. -Stopping the pipeline pauses processing for all sources. To resume, open **Pipeline actions**, select **Start pipeline**, and confirm with **Start pipeline**. Use the [source actions](#stop-and-start-source) to control an individual source separately. Stopping all sources individually leaves the processor running. See [Billing]({{< relref "/operate/rc/rdi/faq#billing" >}}) for the difference in charges. +Stopping the pipeline pauses processing for all sources. To resume, open **Pipeline actions**, select **Start pipeline**, and confirm with **Start pipeline**. Use the [source actions](#stop-and-start-source) to control an individual source separately. Stopping all sources individually leaves the processor running. See [Billing](/content/operate/rc/rdi/faq.md#billing) for the difference in charges. ## Flush the target database @@ -237,18 +236,17 @@ Flushing permanently deletes **all data** from the target database, including re 1. Open **Pipeline actions** and select **Flush target database**. This action is disabled while the pipeline is running. 1. Check the target database and confirm with **Flush target database**. -To refill the target, follow [How do I reload data after a flush?]({{< relref "/operate/rc/rdi/faq#reload-after-flush" >}}). +To refill the target, follow [How do I reload data after a flush?](/content/operate/rc/rdi/faq.md#reload-after-flush). ## Delete pipeline -{{< note >}} -You cannot delete a pipeline while the target database's Redis Cloud cluster is -in maintenance mode. Wait until maintenance is complete, then delete the -pipeline. -{{< /note >}} +> [!NOTE] +> You cannot delete a pipeline while the target database's Redis Cloud cluster is +> in maintenance mode. Wait until maintenance is complete, then delete the +> pipeline. 1. Return to the workspace list on the **Data Integration** page or your database's **Data Integration** tab. 1. Open the pipeline's actions menu and select **Delete pipeline**. 1. Review and confirm the deletion. -Deleted pipelines cannot be recovered. If the pipeline has an associated setup draft, deletion also removes that draft. You can then [delete the workspace]({{}}) when it is no longer needed. +Deleted pipelines cannot be recovered. If the pipeline has an associated setup draft, deletion also removes that draft. You can then [delete the workspace](/content/operate/rc/rdi/create-workspace.md#delete-workspace) when it is no longer needed.