Back to blog EUC & Cloud Workplace

Azure Virtual Desktop at thousands of users: what actually matters in practice

DN
Drazen Nikolic
| | 11 min read

Azure Virtual Desktop is set up quickly. One host pool, one image, a few session hosts: the first user can sign in after two days. That is exactly why many projects underestimate what comes afterwards.

Problems do not appear during the build. They appear in operations: with 200 simultaneous logons at eight in the morning, with profiles that no longer load, and with cost nobody predicted.

AVD projects rarely fail on technology

In practice, AVD platforms break at four points: profile management, badly sliced host pools, autoscaling rules that do not match real working hours, and the image lifecycle. All four are solvable, but they must be decided before the rollout, not after.

The most expensive mistake is treating an AVD environment like a classic terminal server farm. The concepts look similar, the operating model does not: capacity is elastic, images are disposable, and profiles no longer live on the server where the work happens.

Host pool design: density versus stability

Windows 11 multi-session allows high user density per VM, the main cost lever compared to dedicated Cloud PCs. The only question is how many users per host. The honest answer is that there is no universal number. An office user behaves very differently from a developer with a container runtime or an engineer with a CAD application.

What works: separating user groups with similar load profiles into their own host pools instead of building one large pool for everyone. That costs marginally more administration but prevents a single resource-hungry session from slowing down thirty others. In larger environments, pools per region or site follow, so latency to the user stays acceptable.

The second design mistake is choosing the VM family by price instead of by profile. Too little memory per user causes paging and makes sessions feel sluggish while CPU metrics look harmless. GPU workloads belong in dedicated pools with matching VM sizes.

FSLogix decides the logon time

When users complain about an AVD environment, it is usually about two things: logon takes too long, or settings disappeared. Both almost always lead back to profile management.

FSLogix stores the user profile in a container that is attached at logon. That works well, as long as the underlying storage is fast enough, exclusions are set correctly, and nobody lets the antivirus scan the container. They are the first three things to check when users report slow logons.

Then there is the question of what belongs in the profile at all. A container that grows to 30 GB over months because Teams caches and downloads travel with it extends every logon and every backup. Regular compacting and clear redirections are not fine-tuning, they are basic operations.

Anyone who needs transparency here cannot avoid monitoring: how long does attaching the container actually take? Which profiles are corrupted? Which sessions continue without a container and lose data at the end? Questions like these were the reason several open-source tools for the FSLogix community emerged from this operational practice.

Autoscaling saves money only with the right schedule

The biggest cost advantage of AVD comes from capacity disappearing at night and on weekends. This is also where configuration is most often too cautious: shutting hosts down at 10 pm and starting them at 5 am gives away a significant part of the possible saving.

The opposite is equally true: scaling too aggressively produces complaints. If thirty users sign in while the second host is still booting, the saving is not worth the friction. The answer lies in real usage data (logon times across several weeks, separated by weekday), not in assumptions from a project workshop.

One detail is often overlooked: drain mode and shutdown sequencing. Sessions must not be terminated mid-work just because a scaling rule triggered. Configured properly, users never notice it.

Images: the underestimated operational effort

An AVD image is not a one-time artifact, it is a process. Every month brings Windows updates, application updates and security settings. Done manually, you end up with three images in circulation after six months and nobody knows which one runs where.

What works: a defined build process with traceable steps: base image, optimizations, applications, hardening, sysprep pre-checks, versioning. Whether that runs through Azure Image Builder, a script framework or a pipeline is secondary. What matters is reproducibility and documentation of what is inside the image.

Optimization includes points that look trivial individually and are clearly noticeable in sum: removing unnecessary services and apps, setting Defender exclusions for FSLogix containers correctly, preparing multimedia redirection and RDP Shortpath, adjusting power settings. On a platform with thousands of sessions, exactly this determines perceived speed.

What success is actually measured by

For IT, an AVD platform is an architecture topic. For users it is a single number: the time between click and a usable desktop. If logon is reliably fast, users accept almost any change. If it is not, no architecture diagram helps.

That is why it pays to define before the rollout which values are measured: logon duration broken down by phase, container attach time, session density per host, utilization during peak hours. Without that baseline it is impossible to tell later whether a complaint is an individual case or a structural problem.

At this scale a cautious start pays off: a small pilot with real users and measurements, then a gradual rollout. Migrating everyone first and optimizing afterwards costs considerably more.

Need an independent view on your Azure environment?

Calandor reviews Azure architecture, cost and security and delivers prioritised findings your team can act on: as an architecture review, a second opinion or a cost and migration analysis.

Book an initial call