ResearchPod Summary
Can pre-trained diffusion models be adapted for multi-task dense prediction without the computational overhead of parameter-heavy adapters, experts, or learnable task tokens? The authors investigate whether the native, fixed timestep embedding in one-step diffusion models can serve as an endogenous signal to steer the model toward different task-specific manifolds.
The authors introduce Multi-task Unified eStimation via timestep Embedding (MUSE). Instead of adding external modules, MUSE treats the fixed timestep embedding as a categorical semantic switch. During training, different tasks (e.g., depth estimation, surface normal estimation) are assigned unique, discrete timestep values. This conditions the shared U-Net or DiT backbone to activate task-specific computational pathways. The authors interpret this mechanism through Manifold Decoupling, where the orthogonal timestep keys guide the model to project input features onto distinct, non-conflicting task manifolds in the latent space.
MUSE achieves highly competitive performance across 10 dense prediction benchmarks while remaining parameter-free and efficient. The authors demonstrate that the timestep steering mechanism is robust and generalizes across both U-Net and Diffusion Transformer (DiT) architectures. Visualizations of the latent space confirm that features conditioned by different timesteps form clearly separated clusters, validating the hypothesis that the model learns to decouple tasks geometrically rather than through explicit architectural partitioning.
This work provides a minimalist, efficient path toward generalist vision models. By unlocking the latent potential already present in existing generation infrastructure, MUSE eliminates the need for complex multi-task architectures, offering a scalable and resource-efficient solution for unified visual perception.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.