Alibaba DataWorks

AlibabaAnalyticsFree tier available

End-to-end data development platform over MaxCompute, EMR, and Hologres with visual and SQL-based task authoring, scheduled pipelines, data quality rules, data lineage, and a built-in business-glossary catalog

Jurisdictional exposure

Provider HQ
CNHangzhou, China

Subject to PIPL, DSL, CSL

Region locations
APACCNEUUKUSOther26 regions across 6 jurisdictions
Sovereign option
Yes — 11 sovereign-flagged regions available

Attributes

Visual Editor
Yes

Sub-services (4)

Data Studio

Visual and SQL-based task authoring against MaxCompute / EMR / Hologres

Scheduler

Cron / DAG scheduling with dependency resolution and retry policies

Data Quality

Rule-based validation and anomaly detection for pipeline output tables

Data Map

Business glossary, lineage, and discovery portal over DataWorks sources

Compliance & Certifications

This service is attested for the following frameworks. Always verify with the provider before relying on a specific compliance posture.

Where this runs

26 regions
15 countries
11sovereign
Sovereign regions (11)
  • China (Hangzhou) · HangzhouAlibaba Cloud China
  • China (Beijing) · BeijingAlibaba Cloud China
  • China (Shanghai) · ShanghaiAlibaba Cloud China
  • China (Shenzhen) · ShenzhenAlibaba Cloud China
  • China (Chengdu) · ChengduAlibaba Cloud China
  • China (Zhangjiakou) · ZhangjiakouAlibaba Cloud China
  • China (Hohhot) · HohhotAlibaba Cloud China
  • China (Qingdao) · QingdaoAlibaba Cloud China
  • China (Heyuan) · HeyuanAlibaba Cloud China
  • China (Ulanqab) · UlanqabAlibaba Cloud China
  • China (Wuhan) · WuhanAlibaba Cloud China
Commercial regions (15)

Europe (2)

  • Frankfurt
  • London

North America (2)

  • Silicon Valley
  • Virginia

Asia (9)

  • Hong Kong
  • Mumbai
  • Jakarta
  • Tokyo
  • Kuala Lumpur
  • Manila
  • Singapore
  • Seoul
  • Bangkok

Oceania (1)

  • Sydney

Middle East (1)

  • Dubai

Tags

Equivalent services on other platforms

AWS GlueAWS

Serverless data integration platform with visual ETL authoring via Glue Studio, a Hive-compatible Data Catalog, automatic crawlers, DataBrew visual prep, and Glue Data Quality for declarative rule-based validation across Spark and Python jobs

Azure Data FactoryAzure

Managed ETL and ELT service for data integration at scale with 100+ connectors, visual pipeline designer, mapping data flows, and triggers for event-driven orchestration

Databricks Delta Live TablesDatabricks

Declarative ETL framework for streaming and batch pipelines on the lakehouse — define tables in SQL or Python, DLT handles dependency graph, retries, data quality, and observability

DataflowGCP

Unified stream and batch data processing service running Apache Beam pipelines with autoscaling, exactly-once semantics, and native sinks to BigQuery and Cloud Storage

Dataplex Universal CatalogGCP

Intelligent data fabric and catalog that unifies distributed data across Cloud Storage, BigQuery, and third-party lakes with automatic discovery, quality scoring, data lineage, attribute-based access control, and generative AI-powered metadata enrichment

DatastreamGCP

Serverless change-data-capture (CDC) and replication service that continuously streams change events from MySQL, PostgreSQL, Oracle, and SQL Server sources into BigQuery, Cloud Storage, Spanner, and Cloud SQL with minimal source-side impact

Cloud Data FusionGCP

Fully managed enterprise data integration service built on open-source CDAP, with a visual drag-and-drop pipeline builder, 150+ pre-built connectors and transformations, and GitOps-friendly pipeline export for hybrid ETL across Google Cloud and third-party data sources

Pricing

Pricing model:subscription