An Azure service that provides serverless Kubernetes, an integrated continuous integration and continuous delivery experience, and enterprise-grade security and governance.
Hello Marcin van de Ven,
Welcome to the Microsoft Q&A and thank you for posting your questions here.
I understand that Cluster performance metrics not shown for you in the Portal.
Regarding your explanations. The core problem is not AKS workload instability first, it is a user-specific monitoring data access/rendering issue, plus a separate workload configuration issue shown by “CPU/RAM limit is not set up.”
To fix the issue, grant yourself access to all monitoring data planes used by AKS Monitor/Insights, not only the AKS resource IAM blade.
- Grant the affected user Monitoring Reader or Monitoring Contributor on the AKS cluster, the Log Analytics workspace, the Azure Monitor workspace used by Managed Prometheus, and related Data Collection Rules. - https://learn.microsoft.com/en-us/azure/azure-monitor/containers/kubernetes-monitoring-enable, https://learn.microsoft.com/en-us/azure/azure-monitor/metrics/metrics-troubleshoot
- Grant Kubernetes API read access separately if the cluster uses Microsoft Entra/Azure RBAC for Kubernetes authorization, for example Azure Kubernetes Service RBAC Reader at cluster or namespace scope. - https://docs.azure.cn/en-us/aks/concepts-identity, https://learn.microsoft.com/en-us/azure/aks/entra-id-authorization
- Fix the “CPU/RAM limit is not set up” message by adding resource requests and limits to the pod/container manifests. - https://learn.microsoft.com/en-us/azure/aks/developer-best-practices-resource-management, not just by the permission. - https://learn.microsoft.com/en-us/azure/aks/kubernetes-portal
- If the Monitor graphs still show “An unknown error occurred,” check browser/ad-blocking and Azure Monitor workspace networking/private-link access. - https://learn.microsoft.com/en-us/azure/azure-monitor/metrics/metrics-troubleshoot, https://docs.azure.cn/en-us/azure-monitor/containers/container-insights-experience-v2
Therefore, even if AKS resource IAM appears equal. AKS pod/resource views depend on a combination of AKS resource IAM, Kubernetes API authorization, Log Analytics workspace access, Azure Monitor workspace/Prometheus access, and sometimes browser/network access to Monitor endpoints. You will mneed to have complete effective access to the monitoring data plane used by the AKS portal experience. Fix those permissions directly, then separately add CPU/memory requests and limits to the workloads. This restores the portal metrics and explains the “CPU/RAM limit is not set up” message.
I hope this is helpful! Do not hesitate to let me know if you have any other questions, steps or clarifications.
Please don't forget to close up the thread here by upvoting and accept it as an answer if it is helpful.