Fetching the paper…
Reading the bibliography…
We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions.
B. K. P. Horn and B. G. Schunck, “Determining optical flow,” Artificial Intelligence , vol. 17, no. 1, pp. 185–203, 1981
1981
Earlier work this paper cites.
D. S. Turaga, Y. Chen, and J. Caviedes, “No reference psnr estimation for compressed pictures,” Signal Processing: Image Communication , vol. 19, no. 2, pp. 173–184, 2004
2004
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, U-Net: Convolutional Networks for Biomedical Image Segmentation , ser. Lecture Notes in Computer Science, 2015, book section Chapter 28, pp. 234–241
2015
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” 2016 Ieee Conference on Computer Vision and Pattern Recognition (Cvpr) , pp. 2818–2826, 2016
2016
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , ser. Series GANs trained by a two time-scale update rule converge to a local nash equilibrium. Curran Associates Inc., 2017 Published, Conference Paper, p. 6629–6640
2017
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=St1giarCHLP
2021
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in Proceedings of the 39th International Conference on Machine Learning , ser. Series BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, C. Kamalika, J. Stefanie, S. Le, S. Csaba, N. Gang, and S. Sivan, Eds., vol. 162. PMLR, 2022 Published, Conference Paper, pp. 12 888–12 900
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Bar-Tal, D. Ofri-Amar, R. Fridman, Y. Kasten, and T. Dekel, Text2LIVE: Text-Driven Layered Image and Video Editing , ser. Lecture Notes in Computer Science, 2022, book section Chapter 41, pp. 707–723
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Closest in time.