Fetching the paper…

RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models · Around