An Azure service that is used to provision Windows and Linux virtual machines.
Hello @Luis Dinis
As discussed over the teams call pls follow this steps.
Step 1 — Protect your data first (most critical given disk swap failed)
Before anything else, snapshot the disks so you have a safe fallback.
- In the Azure portal, go to your VM → Disks
- Click on the OS disk (and each data disk) → Create snapshot
- Give the snapshot a name and choose the same resource group
- Repeat for all data disks containing user/personal data
This ensures you can recover data no matter what changes you make. This step is specifically called out because the disk swap didn't work — your data is still on the original disk, and you must protect it before proceeding.
Step 2 — Verify VM and Bastion health
Check the VM:
- Portal → go to the affected VM → Overview
- Confirm Status = Running
- Go to Boot diagnostics → Screenshot — confirm the OS shows a login screen, not a boot error or blank screen
- Check Activity log for any recent failures, restarts, or platform events
Check Bastion:
- Portal → search for Bastion → open your Bastion resource
- Verify Provisioning state = Succeeded
- Open Diagnose and solve problems — look for any active health alerts
- Try connecting from a different browser (Chrome/Edge) and a different network — this rules out client-side proxy or browser extension interference
Step 3 — Run Connection Troubleshoot (most diagnostic step)
This is the single most effective tool and all three AI responses agreed on it.
- Portal → Network Watcher → Connection troubleshoot
- Fill in:
- Source type: Virtual machine (pick a healthy VM in the same VNet) or use Bastion as source
- Destination type: Virtual machine → select the affected VM
- Protocol: TCP
- Port: 3389 (Windows RDP) or 22 (Linux SSH)
- Click Check
- Interpret results:
- Reachable → the network path is fine; the issue is OS-level (guest firewall, RDP/SSH service, credentials)
- Unreachable → the issue is network: NSG, route table, or VM OS crash/hang
- Reachable → the network path is fine; the issue is OS-level (guest firewall, RDP/SSH service, credentials)
- Protocol: TCP
- Destination type: Virtual machine → select the affected VM
- Source type: Virtual machine (pick a healthy VM in the same VNet) or use Bastion as source
Step 4 — Validate NSG rules
This is the most common silent cause of Bastion failures.
On the AzureBastionSubnet NSG:
- Portal → Virtual networks → your VNet → Subnets → AzureBastionSubnet
- Open the attached NSG
- Verify required inbound rules exist (ports 443, 8080 from GatewayManager and Internet)
- Verify required outbound rules exist (ports 3389/22 to the VM subnet, ports 443/8080 to Internet and AzureCloud)
- Make sure there is no custom deny rule with a lower priority number overriding these
On the VM's NIC/subnet NSG:
- Go to VM → Networking → click the NSG
- Confirm there is an Allow inbound rule for:
- Protocol: TCP
- Port: 3389 (Windows) or 22 (Linux)
- Source: use the tag
VirtualNetworkor specify the Bastion subnet CIDR explicitly
- Confirm no higher-priority Deny rule is blocking the same port
- Source: use the tag
- Port: 3389 (Windows) or 22 (Linux)
- Protocol: TCP
Step 5 — Check for UDRs and DNS conflicts
A force-tunnel route is a very common, hard-to-spot cause of unstable Bastion sessions.
- Portal → VM's subnet → Route table (if any is associated)
- Check for any route with:
- Address prefix:
0.0.0.0/0- Next hop type: Virtual appliance or VPN gateway
- This force-tunnels all traffic through a firewall/NVA that may be dropping the Bastion session
- If using a custom/private DNS zone, verify the VM can resolve standard Azure hostnames, or test connecting by private IP directly in Bastion instead of DNS name
- Next hop type: Virtual appliance or VPN gateway
- Address prefix:
Step 6 — Fix the OS-level service (safe, no data loss)
If the network path is confirmed fine (Step 3 says Reachable), the RDP or SSH service inside the VM may be broken. Fix it without touching the disk using Run Command:
For Windows:
- VM → Operations → Run command → select
EnableRDP - Alternatively, run
SetRDPPortand set port to 3389 - This re-enables RDP and adjusts the Windows Firewall rule automatically
For Linux:
- VM → Run command →
RunShellScript - Enter:
sudo systemctl status sshd— check if the service is active - If stopped:
sudo systemctl restart sshd - Check firewall:
sudo ufw statusorsudo iptables -L
Step 7 — VM recovery actions (if still unresponsive)
- Restart: VM → Overview → Restart — this alone resolves many unstable Bastion states
- Redeploy: VM → Help → Redeploy + reapply — moves the VM to a fresh Azure host node, preserving all data
- If the VM fails to boot (boot diagnostics shows errors), attach the OS disk to a repair VM:
- Create a new VM in the same region/resource group
- VM → Disks → Detach the OS disk (only if VM is deallocated)
- Attach it as a data disk on the repair VM
- Mount and access user data from there
- Attach it as a data disk on the repair VM
- VM → Disks → Detach the OS disk (only if VM is deallocated)
- Create a new VM in the same region/resource group
Step 8 — Isolate: Bastion path vs VM issue
If you're still unsure whether Bastion itself or the VM is the problem:
- Deploy a test VM in the same VNet and subnet
- Try to RDP/SSH from the test VM directly to the affected VM's private IP (peer-to-peer, not through Bastion)
- Interpret:
- Works → the VM is healthy; the fault is in the Bastion path (Bastion config, its NSG, or the subnet)
- Fails → the VM itself or its subnet/NSG is the issue
- Works → the VM is healthy; the fault is in the Bastion path (Bastion config, its NSG, or the subnet)
Since the disk swap didn't fix it, the OS/data is almost certainly fine. The priority order for root causes is:
- NSG rule missing on the VM NIC or AzureBastionSubnet — most common
- Bastion connectivity glitch — restart the VM + try from another browser first
- UDR / force-tunnel route — check the subnet route table
- OS firewall or service crash — use Run Command to fix without data loss
Thanks,
Manish