Kerberos fails on AG Listener — only from cluster nodes, ticket not generated

Giorgio 0 Reputation points
2026-05-22T15:33:45.5866667+00:00

Environment: SQL Server Always On AG on top of FCI, Windows Server cluster, domain-joined nodes.

Problem: Connections to the AG listener always fall back to NTLM, but only when connecting from the cluster nodes themselves. From any other machine on the network, Kerberos works correctly.

Symptoms:

  • sys.dm_exec_connections.auth_scheme returns NTLM on listener connections from cluster nodes
  • klist shows no ticket generated after a SQL connection attempt
  • No Kerberos Event ID 4769 on the Domain Controller during the connection attempt
  • klist get MSSQLSvc/<listener-fqdn>:<port> does return a valid AES-256 ticket when requested manually
  • After a SQL connection, the ticket does not appear in klist — SSPI never requests it

What was already ruled out:

  • DNS resolution — correct
  • Duplicate SPNs (setspn -X -F) — none found
  • DisableLoopbackCheck and BackConnectionHostNames — already set, no effect
  • AD computer account for the listener VCO — enabled
  • Trust relationship on cluster nodes — healthy

Current finding: The ticket can be obtained manually, meaning the SPN is registered and the KDC responds correctly. However, the SQL driver / SSPI never requests the ticket during connection — it goes directly to NTLM without contacting the KDC at all.

Question: Has anyone seen SSPI silently bypass Kerberos on AG listener connections from cluster nodes, specifically in an AG-over-FCI topology? Any way to force the driver to use a specific SPN or disable name canonicalization for this target?NOTE : I suppose the problem begins after April update...because it was running for a long period

SQL Server | Other
SQL Server | Other

Additional SQL Server features and topics not covered by specific categories

0 comments No comments

Answer recommended by moderator
Giorgio 0 Reputation points
2026-05-29T11:40:52.16+00:00

Thanks all... i find the solution! When I installed the cluster in 2024 I've set :

HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Lsa\MSV1_0 BackConnectionHostNames = (REG_MULTI_SZ)

with all VVN of my cluster instance, to increase performance of linked server between multiple SQL server instance.

and it stopped working in April 2026, i suppose due updates to kerberos.

I've remove all entry from the registry key and after reboot all new session are using kerberos auth.

Was this answer helpful?

0 comments No comments

2 additional answers

Sort by: Most helpful
  1. Anonymous
    2026-05-25T08:55:08.1766667+00:00

    Hi @Giorgio ,

    Since klist get works manually and Kerberos is fine from other machines, your SPN and KDC are not the problem.

    The missing 4769 on the DC is the real tell SSPI isn't even trying Kerberos from the cluster nodes, it's bailing before it contacts the KDC. This behavior tightened after the April update (CVE-2022-21920 hardening), cluster nodes are more exposed because of how they resolve the listener name internally.

    Few things to try in order:

    • klist purge on each node, then retest
    • Check %SystemRoot%\System32\drivers\etc\hosts on each node for any listener/VNN entries these sometimes sneak in during cluster setup
    • Run nslookup <listener-fqdn> on each node and make sure what comes back exactly matches what's in setspn -L <sql-service-account>
    • Connect using tcp:<listener-fqdn>,<port> explicitly, avoid short names
    • Test with ServerSPN=MSSQLSvc/<listener-fqdn>:<port> in the connection string

    If that last one fixes it, the name canonicalization on the cluster nodes is your culprit. You can either keep ServerSPN explicitly in the connection string (it's fully supported) or track down the DNS/HOSTS mismatch and fix it properly.

    Thanks,
    Akhil.

    Was this answer helpful?

    0 comments No comments

  2. AI answer

    2026-05-22T15:34:10.26+00:00

    This pattern matches a Kerberos/SPN/SSPI behavior change where, after recent security updates, Kerberos is not even attempted if the SPN or name resolution doesn’t meet stricter rules. Because:

    • Kerberos works from other machines to the listener.
    • klist get MSSQLSvc/<listener-fqdn>:<port> succeeds from the nodes (SPN and KDC are fine).
    • But during an actual SQL connection from the nodes, SSPI never requests a ticket and falls straight to NTLM.

    The key is that the SQL client builds its own SPN from the resolved name and may reject it under newer rules, so it never calls the KDC.

    Actionable checks and fixes based on the provided material:

    1. Validate the exact SPN format used by the client
      • For TCP connections, the driver auto‑generates MSSQLSvc/FQDN:[<port>].
      • If WinSock name resolution on the cluster nodes returns a different FQDN (extra DNS suffix, wrong alias, HOSTS override, etc.), the client may form an SPN that is considered malformed or doesn’t exist and then skip Kerberos.
      • On each cluster node, run name resolution tests against the listener:
        • ping <listener-fqdn>
        • nslookup <listener-fqdn>
      • Ensure the resolved FQDN exactly matches the SPN that is registered for the listener and the SQL Server service account.
    2. Check for HOSTS file interference on the nodes
      • An incorrect HOSTS entry can cause the client to generate an unexpected SPN and fall back to NTLM.
      • On each node, review %SystemRoot%\System32\drivers\etc\hosts and remove any entries for the listener or its IP.
      • This is explicitly called out as a cause of NTLM fallback and unexpected SPN generation.
    3. Re‑verify SPN registration for AG over FCI
      • For Kerberos on an Always On listener, all instances hosting replicas must use the same SQL Server service account.
      • The domain admin must register the SPN on that service account for the listener VNN (not the node names), in the MSSQLSvc/<listener-fqdn>:<port> form.
      • If the SPN is registered on any account other than the SQL Server service account, Kerberos fails.
      • After any service account change, SPNs must be re‑registered.
    4. Consider malformed/stricter SPN validation after updates
      • Recent security hardening (for example, protections like those described for CVE‑2022‑21920) can block NTLM fallback when Kerberos fails for malformed or non‑existent SPNs.
      • Common causes:
        • SPN format not matching what the API expects (wrong host part, port on the wrong segment, non‑DNS name where DNS is required).
        • SPN missing entirely for the exact 3‑part name the client is using.
      • Confirm that the SPN used by the client is considered legitimate per the documented patterns; if necessary, register the exact SPN variant that the client is generating to the correct SQL Server service account.
    5. Confirm no duplicate or wrong‑account SPNs
      • Duplicate SPNs or SPNs mapped to the wrong account cause Kerberos failures and can force NTLM.
      • Use setspn -X -F (already done) and also explicitly list SPNs on the SQL Server service account and any related computer accounts to ensure:
        • No duplicates.
        • No MSSQLSvc/<listener-fqdn> SPN exists on a different account than the SQL Server service account.
    6. Check for expired or stale Kerberos tickets on the nodes
      • If the node is holding expired or stale tickets, Kerberos can fail silently and the client may fall back to NTLM.
      • On each node:
        • Run klist purge, then log off/on or restart.
        • Re‑test the listener connection and check sys.dm_exec_connections.auth_scheme again.
    7. Inspect for broader authentication/SSPI issues on the nodes
      • Review the system and security event logs on the nodes around the time of connection attempts for Kerberos, LSA, or Netlogon errors.
      • If there are errors like SSPI handshake failed with specific Win32 codes (for example, 0x80090311 or 0x80090304), follow the guidance to check domain controller reachability and AD health from those nodes.
    8. Forcing SPN / disabling canonicalization
      • SQL Server client drivers support user‑specified SPNs via connection attributes. When provided, the driver uses that SPN directly instead of generating one from DNS.
      • Use a connection string that explicitly sets the SPN (for example, via the driver’s “server SPN” or “Authentication”/“ServerSPN” attributes, depending on provider) in the form:
        • MSSQLSvc/<listener-fqdn>:<port>
      • This bypasses name canonicalization issues and forces SSPI to request that exact SPN.
      • If this change restores Kerberos from the cluster nodes, it confirms that the auto‑generated SPN or name canonicalization on those nodes is the root cause.
    9. AG‑over‑FCI specifics
      • In AG over FCI, ensure:
        • All replicas (FCI instances) use the same domain service account.
        • SPNs exist for both the AG listener VNN and any database mirroring endpoints if Kerberos is required for endpoint communication.
      • If the April update introduced stricter checks, any previously “forgiving” SPN or DNS misalignment on the nodes can now cause SSPI to skip Kerberos only from those nodes.

    In practice, start by:

    1. Purging tickets on the nodes (klist purge) and re‑testing.
    2. Removing any HOSTS entries and confirming DNS returns the exact listener FQDN used in the SPN.
    3. Testing a connection that explicitly specifies MSSQLSvc/<listener-fqdn>:<port> as the SPN in the connection string.
    4. Re‑validating that the listener SPN is registered only on the SQL Server service account and matches the listener FQDN and port.

    If Kerberos then works from the nodes, the issue is confirmed as SPN/name canonicalization behavior tightened by recent updates.


    References:

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.