We are using Azure Data Factory Copy Activity to extract ZIP files and write the extracted files to Azure Data Lake Storage Gen2.
The Copy Activity fails for some files/directories inside the ZIP with the following error:

The ZIP contains directories/file names with non-ASCII characters. Other ZIP files containing non-English characters, including Portuguese characters, are successfully extracted by the same ADF Copy Activity.
Therefore, this does not appear to be a general restriction against non-English characters.
One observation is that the failing path contains strings such as NgÃ..., Tñ, Ä..., etc. These appear similar to UTF-8 characters being decoded using an incorrect code page (UTF-8 mojibake).
Questions
- Does Azure Data Factory Copy Activity use the .NET ZIP extraction library when decompressing ZIP files?
- Does ADF Copy Activity correctly support UTF-8/Unicode ZIP entry names?
- How does ADF determine the encoding of ZIP entry filenames when extracting a ZIP?
- Does ADF respect the ZIP General Purpose Bit Flag indicating UTF-8 encoding for entry names?
- Is there currently any Copy Activity setting to explicitly specify UTF-8 as the encoding for ZIP entry names?
- Could an incorrectly encoded ZIP entry name result in ADF generating an invalid ADLS Gen2 path and therefore returning HTTP 400 BadRequest?
- Is there an official Microsoft guideline specifying which Unicode characters are supported/not supported in ADLS Gen2 file and directory names?
We would also like to understand whether this issue is caused by:
A. An illegal Unicode character in the original filename, or
B. Incorrect decoding/encoding of the ZIP entry filename during ADF extraction.
The same directory/file name can potentially be created directly in ADLS Gen2, but the failure occurs specifically when ADF extracts the ZIP.
Any guidance on how to validate the ZIP filename encoding or identify the exact Unicode code point that ADF is rejecting would be appreciated.
Environment
- Azure Data Factory
- Copy Activity
- Source: ZIP file
- Sink: Azure Data Lake Storage Gen2
- Compression/extraction: ZipDeflateThe ZIP contains directories/file names with non-ASCII characters. Other ZIP files containing non-English characters, including Portuguese characters, are successfully extracted by the same ADF Copy Activity. Therefore, this does not appear to be a general restriction against non-English characters. One observation is that the failing path contains strings such as
NgÃ..., Tñ, Ä..., etc. These appear similar to UTF-8 characters being decoded using an incorrect code page (UTF-8 mojibake). Questions
- Does Azure Data Factory Copy Activity use the .NET ZIP extraction library when decompressing ZIP files?
- Does ADF Copy Activity correctly support UTF-8/Unicode ZIP entry names?
- How does ADF determine the encoding of ZIP entry filenames when extracting a ZIP?
- Does ADF respect the ZIP General Purpose Bit Flag indicating UTF-8 encoding for entry names?
- Is there currently any Copy Activity setting to explicitly specify UTF-8 as the encoding for ZIP entry names?
- Could an incorrectly encoded ZIP entry name result in ADF generating an invalid ADLS Gen2 path and therefore returning HTTP 400 BadRequest?
- Is there an official Microsoft guideline specifying which Unicode characters are supported/not supported in ADLS Gen2 file and directory names?
We would also like to understand whether this issue is caused by: A. An illegal Unicode character in the original filename, or B. Incorrect decoding/encoding of the ZIP entry filename during ADF extraction. The same directory/file name can potentially be created directly in ADLS Gen2, but the failure occurs specifically when ADF extracts the ZIP. Any guidance on how to validate the ZIP filename encoding or identify the exact Unicode code point that ADF is rejecting would be appreciated. Environment
- Azure Data Factory
- Copy Activity
- Source: ZIP file
- Sink: Azure Data Lake Storage Gen2
- Compression/extraction: ZipDeflate