← 返回时间线

paper

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

arXiv ↗
ID
2609.07414
分类
首次捕获
2026-09-10
状态
unread
作者
Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang
信号
🔥 8

信号历史

  • 2026-09-10HF Daily Papers · 🔥8
暂无信号数据

摘要

Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.

我的笔记

还没有笔记。

在 GitHub 上写笔记 ↗(新建 content/notes/2609.07414.md,PR 合并后本页自动更新)