Azure Data Factory Copy Activity fails extracting ZIP files to ADLS Gen2 with BadRequest for Unicode/Vietnamese filenames — possible ZIP filename encoding issue

Balasuriya R 0 Reputation points
2026-09-17T04:19:23.85+00:00

We are using Azure Data Factory Copy Activity to extract ZIP files and write the extracted files to Azure Data Lake Storage Gen2.

The Copy Activity fails for some files/directories inside the ZIP with the following error:
User's image

The ZIP contains directories/file names with non-ASCII characters. Other ZIP files containing non-English characters, including Portuguese characters, are successfully extracted by the same ADF Copy Activity.

Therefore, this does not appear to be a general restriction against non-English characters.

One observation is that the failing path contains strings such as NgÃ..., Tñ, Ä..., etc. These appear similar to UTF-8 characters being decoded using an incorrect code page (UTF-8 mojibake).

Questions

  1. Does Azure Data Factory Copy Activity use the .NET ZIP extraction library when decompressing ZIP files?
  2. Does ADF Copy Activity correctly support UTF-8/Unicode ZIP entry names?
  3. How does ADF determine the encoding of ZIP entry filenames when extracting a ZIP?
  4. Does ADF respect the ZIP General Purpose Bit Flag indicating UTF-8 encoding for entry names?
  5. Is there currently any Copy Activity setting to explicitly specify UTF-8 as the encoding for ZIP entry names?
  6. Could an incorrectly encoded ZIP entry name result in ADF generating an invalid ADLS Gen2 path and therefore returning HTTP 400 BadRequest?
  7. Is there an official Microsoft guideline specifying which Unicode characters are supported/not supported in ADLS Gen2 file and directory names?

We would also like to understand whether this issue is caused by:

A. An illegal Unicode character in the original filename, or

B. Incorrect decoding/encoding of the ZIP entry filename during ADF extraction.

The same directory/file name can potentially be created directly in ADLS Gen2, but the failure occurs specifically when ADF extracts the ZIP.

Any guidance on how to validate the ZIP filename encoding or identify the exact Unicode code point that ADF is rejecting would be appreciated.

Environment

  • Azure Data Factory
  • Copy Activity
  • Source: ZIP file
  • Sink: Azure Data Lake Storage Gen2
  • Compression/extraction: ZipDeflateThe ZIP contains directories/file names with non-ASCII characters. Other ZIP files containing non-English characters, including Portuguese characters, are successfully extracted by the same ADF Copy Activity. Therefore, this does not appear to be a general restriction against non-English characters. One observation is that the failing path contains strings such as NgÃ..., Tñ, Ä..., etc. These appear similar to UTF-8 characters being decoded using an incorrect code page (UTF-8 mojibake). Questions
    1. Does Azure Data Factory Copy Activity use the .NET ZIP extraction library when decompressing ZIP files?
    2. Does ADF Copy Activity correctly support UTF-8/Unicode ZIP entry names?
    3. How does ADF determine the encoding of ZIP entry filenames when extracting a ZIP?
    4. Does ADF respect the ZIP General Purpose Bit Flag indicating UTF-8 encoding for entry names?
    5. Is there currently any Copy Activity setting to explicitly specify UTF-8 as the encoding for ZIP entry names?
    6. Could an incorrectly encoded ZIP entry name result in ADF generating an invalid ADLS Gen2 path and therefore returning HTTP 400 BadRequest?
    7. Is there an official Microsoft guideline specifying which Unicode characters are supported/not supported in ADLS Gen2 file and directory names?
    We would also like to understand whether this issue is caused by: A. An illegal Unicode character in the original filename, or B. Incorrect decoding/encoding of the ZIP entry filename during ADF extraction. The same directory/file name can potentially be created directly in ADLS Gen2, but the failure occurs specifically when ADF extracts the ZIP. Any guidance on how to validate the ZIP filename encoding or identify the exact Unicode code point that ADF is rejecting would be appreciated. Environment
    • Azure Data Factory
    • Copy Activity
    • Source: ZIP file
    • Sink: Azure Data Lake Storage Gen2
    • Compression/extraction: ZipDeflate
Azure Data Factory
Azure Data Factory

An Azure service for ingesting, preparing, and transforming data at scale.

0 comments No comments

1 answer

Sort by: Newest
  1. Aditya Singh Rathore 195 Reputation points
    2026-09-21T11:54:39.9+00:00

    Hi @Balasuriya R ,

    I don't think ADF is mis-decoding your ZIP. In your error path, Chụp hình trưng bày is fine, but Ngô Gia Tá»± , P. Ä?ằng Lâm is UTF-8 that was read as Latin-1 and saved again. One entry name gets decoded one way, so if ADF had picked the wrong encoding, both parts would be garbled. That suggests the double-encoding is already in the ZIP's names, which fits your option A better than B.

    For the 400, the box after "Ä" is probably U+0090, a control character. The storage naming docs say control characters aren't allowed in blob names, and a name that breaks the rules returns 400 Bad Request. The file name also has literal % sequences like %DC, which are worth testing separately. microsoft

    On your questions 1 to 5: I couldn't find public ADF docs on ZIP entry-name handling, and I'm not aware of a setting for it. If ADF uses .NET's ZipArchive with defaults, UTF-8 is used when the language encoding flag is set, and the system default code page is used when it isn't. I can't confirm what ADF does. Microsoft Learn

    To check, list the entries with Python's zipfile and print each one's UTF-8 flag and any control characters. Then try creating that exact path directly in ADLS with the real U+0090 in it. Copying it from the error text may drop the character. If ADLS also returns 400, the character is the cause.

    Fix the names where the folders are generated, or unzip outside ADF and clean them first. Hope this helps :)

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.