An Azure service for ingesting, preparing, and transforming data at scale.
Hi @Nhan, Phuong ,
Thank you for reaching out to the Microsoft Q&A.
The issue is likely related to processing large XML data rather than the ADF mapping itself. Since the load works for small datasets but hangs when processing more than 2,000 records, consider the following:
- Check the execution time of the
XMLSERIALIZEquery directly in DB2. - Enable source partitioning and parallel copy in ADF.
- Increase DIUs and parallel copy settings.
- Validate whether some XML records are significantly larger than others.
- As a test, load the data to CSV first instead of Parquet to determine if the bottleneck is in Parquet conversion.
Please share:
- XML column size (average/max)
- Azure IR or Self-Hosted IR
- Copy activity monitoring metrics
This will help determine whether the bottleneck is in DB2 query execution, network transfer, or Parquet generation.