TY - GEN
T1 - IDiffPose
T2 - 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
AU - Wicaksono, Nugroho
AU - Muflikhah, Lailil
AU - Yudistira, Novanto
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2025/12/6
Y1 - 2025/12/6
N2 - Recent diffusion-graph frameworks such as DiffPose have advanced monocular 2D-to-3D pose lifting by coupling graph convolutional networks (GCNs) with denoising diffusion models. However, their depth is still bounded by the number of explicit attention-GCN blocks, limiting the receptive field under a fixed parameter budget. We propose IDiffPose, short for Implicit Diffusion Pose, an equilibrium substitution that replaces a selected deep block with a weight-tied operator whose representation is computed via a small number of unrolled fixed-point iterations. The training objective remains standard DDPM-style denoising, inference uses a deterministic DDIM sampler, and the forward process injects heteroscedastic Gaussian noise scaled by 2D heat-map uncertainty. Because the same weights are re-used across iterations, the model attains large effective depth with a compact parameter footprint; when non-equilibrium blocks are frozen, the number of trainable parameters drops substantially. On Human3.6M, substituting only the middle block improves MPJPE from 31.55 to 31.10 mm and P-MPJPE from 24.72 to 24.47 mm with 167,521 trainable parameters (vs. 1,025,674 explicit), while making all blocks implicit reaches 30.90/24.51 mm. These results show equilibrium substitution is a practical drop-in upgrade for diffusion-based pose lifting, yielding consistent accuracy with favorable efficiency controls.
AB - Recent diffusion-graph frameworks such as DiffPose have advanced monocular 2D-to-3D pose lifting by coupling graph convolutional networks (GCNs) with denoising diffusion models. However, their depth is still bounded by the number of explicit attention-GCN blocks, limiting the receptive field under a fixed parameter budget. We propose IDiffPose, short for Implicit Diffusion Pose, an equilibrium substitution that replaces a selected deep block with a weight-tied operator whose representation is computed via a small number of unrolled fixed-point iterations. The training objective remains standard DDPM-style denoising, inference uses a deterministic DDIM sampler, and the forward process injects heteroscedastic Gaussian noise scaled by 2D heat-map uncertainty. Because the same weights are re-used across iterations, the model attains large effective depth with a compact parameter footprint; when non-equilibrium blocks are frozen, the number of trainable parameters drops substantially. On Human3.6M, substituting only the middle block improves MPJPE from 31.55 to 31.10 mm and P-MPJPE from 24.72 to 24.47 mm with 167,521 trainable parameters (vs. 1,025,674 explicit), while making all blocks implicit reaches 30.90/24.51 mm. These results show equilibrium substitution is a practical drop-in upgrade for diffusion-based pose lifting, yielding consistent accuracy with favorable efficiency controls.
KW - 3-D human pose estimation
KW - deep equilibrium networks
KW - diffusion models
KW - graph neural networks
UR - https://www.scopus.com/pages/publications/105025105910
U2 - 10.1145/3743093.3771064
DO - 10.1145/3743093.3771064
M3 - Conference contribution
AN - SCOPUS:105025105910
T3 - Proceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
BT - Proceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
A2 - Chua, Tat-Seng
A2 - Wong, Lai-Kuan
A2 - Chan, Chee Seng
A2 - Tang, Jinhui
A2 - Ngo, Chong-Wah
A2 - Schoeffmann, Klaus
A2 - Liu, Jiaying
A2 - Ho, Yo-Sung
PB - Association for Computing Machinery, Inc
Y2 - 9 December 2025 through 12 December 2025
ER -