This media is not supported in your browser
VIEW IN TELEGRAM
πMagicMakeup Transferπ
πMakeup-transfer applies the reference makeup to the source face while preserving the source identity. Authors: Zhejiang University & vivo BlueImage Lab. Repo for non commercialπ
πReview https://t.ly/JYpCr
πPaper harxiv.org/pdf/2607.20924
πProject vivocameraresearch.github.io/magicmakeup
πRepo github.com/vivoCameraResearch/Magic-Makeup
πMakeup-transfer applies the reference makeup to the source face while preserving the source identity. Authors: Zhejiang University & vivo BlueImage Lab. Repo for non commercialπ
πReview https://t.ly/JYpCr
πPaper harxiv.org/pdf/2607.20924
πProject vivocameraresearch.github.io/magicmakeup
πRepo github.com/vivoCameraResearch/Magic-Makeup
β€4π2π₯1
This media is not supported in your browser
VIEW IN TELEGRAM
πMicroZoom at Extreme Scaleπ
πMicroZoom by UWA synthesizes gigapixel-resolution images grounded in consumer-grade microscope close-ups at magnification levels up to 350Γ. Impressive. Repo under MITπ
πReview https://t.ly/hgJD7
πPaper https://arxiv.org/pdf/2607.24729
πProject https://microzoom-sr.github.io/
πRepo github.com/MicroZoom-SR/MicroZoom-SR.github.io
πMicroZoom by UWA synthesizes gigapixel-resolution images grounded in consumer-grade microscope close-ups at magnification levels up to 350Γ. Impressive. Repo under MITπ
πReview https://t.ly/hgJD7
πPaper https://arxiv.org/pdf/2607.24729
πProject https://microzoom-sr.github.io/
πRepo github.com/MicroZoom-SR/MicroZoom-SR.github.io
π€―5β€3π2π₯1
This media is not supported in your browser
VIEW IN TELEGRAM
πΏ Dawn of Generative Cinematography πΏ
π© The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.
π Meanwhile, AI research is heading in the exact opposite direction.
π© A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.
πMore: https://t.ly/g-qUh
πPaper arxiv.org/pdf/2607.24591
πProject yixuanli98.github.io/cameraanything/
πRepo github.com/yixuanli98/CameraAnything
π© The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.
π Meanwhile, AI research is heading in the exact opposite direction.
π© A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.
πMore: https://t.ly/g-qUh
πPaper arxiv.org/pdf/2607.24591
πProject yixuanli98.github.io/cameraanything/
πRepo github.com/yixuanli98/CameraAnything
π₯7β€5π1π1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯Decoder-only Any-to-Any Modelπ₯
πMODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under Apacheπ
πReview https://t.ly/-2QKT
πPaper https://lnkd.in/dhfBQGhB
πProject https://lnkd.in/dPD_ECXk
πRepo https://lnkd.in/dbDHw24u
πMODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under Apacheπ
πReview https://t.ly/-2QKT
πPaper https://lnkd.in/dhfBQGhB
πProject https://lnkd.in/dPD_ECXk
πRepo https://lnkd.in/dbDHw24u
π₯5β€3π1πΎ1
This media is not supported in your browser
VIEW IN TELEGRAM
π Dual-branch ID-Trackingπ
πTIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT licenseπ
πReview https://t.ly/WEDeY
πPaper https://arxiv.org/pdf/2607.26412
πProject https://vranlee.github.io/TIDE/
πRepo https://github.com/vranlee/TIDE
πTIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT licenseπ
πReview https://t.ly/WEDeY
πPaper https://arxiv.org/pdf/2607.26412
πProject https://vranlee.github.io/TIDE/
πRepo https://github.com/vranlee/TIDE
π₯9β€7π1π1
π₯Unified Points n' Linesπ₯
πETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under Apacheπ
πReview https://lnkd.in/p/eW8j5JZj
πPaper https://arxiv.org/pdf/2608.19894
πRepo https://github.com/francois141/upal
πETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under Apacheπ
πReview https://lnkd.in/p/eW8j5JZj
πPaper https://arxiv.org/pdf/2608.19894
πRepo https://github.com/francois141/upal
β€11π₯7π3πΎ2π1
This media is not supported in your browser
VIEW IN TELEGRAM
π Anyone in 4D is out π
π4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0π
πReview https://lnkd.in/p/ec4dzGvb
πPaper https://arxiv.org/pdf/2608.20335
πProject https://4danyone.github.io
πRepo github.com/ant-research/4DAnyone
π4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0π
πReview https://lnkd.in/p/ec4dzGvb
πPaper https://arxiv.org/pdf/2608.20335
πProject https://4danyone.github.io
πRepo github.com/ant-research/4DAnyone
π₯7β€4π1π1
This media is not supported in your browser
VIEW IN TELEGRAM
πβπ©Remesh-Aware Mesh Deformationπβπ©
πRADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MITπ
πReview https://lnkd.in/p/eK6FZv9c
πPaper https://arxiv.org/pdf/2608.17182
πProject https://threedle.github.io/radmesh/
πRepo https://github.com/threedle/radmesh/
πRADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MITπ
πReview https://lnkd.in/p/eK6FZv9c
πPaper https://arxiv.org/pdf/2608.17182
πProject https://threedle.github.io/radmesh/
πRepo https://github.com/threedle/radmesh/
β€5π₯3π1π1
This media is not supported in your browser
VIEW IN TELEGRAM
π¦Unified Segmentation n' Retrievalπ¦
πFoundYou gets an example of your object and it segments the same physical instance in a new image or retrieve it from a large gallery with ONE super-compact model. Repo/demo availableπ
πReview https://lnkd.in/p/ex2qnKHW
πPaper arxiv.org/pdf/2608.29917
πProject https://lnkd.in/eNEUB_nV
πRepo https://lnkd.in/eRyDY6Ue
πFoundYou gets an example of your object and it segments the same physical instance in a new image or retrieve it from a large gallery with ONE super-compact model. Repo/demo availableπ
πReview https://lnkd.in/p/ex2qnKHW
πPaper arxiv.org/pdf/2608.29917
πProject https://lnkd.in/eNEUB_nV
πRepo https://lnkd.in/eRyDY6Ue
π₯10β€4π1
This media is not supported in your browser
VIEW IN TELEGRAM
πVision Weight Estimationπ
πDoppio is a novel video dataset capturing video of falling ground coffee, paired with precise, per-frame ground-truth weight measurements: OCR readings are extracted from the display, smoothed and time-lag compensated, and paired with per-frame weight annotations. Repo to be released under Apacheπ
πReview https://www.linkedin.com/posts/visionarynet_computer-vision-weight-estimation-activity-7501900695457951744-ArBO
πPaper https://lnkd.in/eHuy87SX
πProject https://lnkd.in/e9g9zeK3
πRepo https://lnkd.in/emUePTiq
πDoppio is a novel video dataset capturing video of falling ground coffee, paired with precise, per-frame ground-truth weight measurements: OCR readings are extracted from the display, smoothed and time-lag compensated, and paired with per-frame weight annotations. Repo to be released under Apacheπ
πReview https://www.linkedin.com/posts/visionarynet_computer-vision-weight-estimation-activity-7501900695457951744-ArBO
πPaper https://lnkd.in/eHuy87SX
πProject https://lnkd.in/e9g9zeK3
πRepo https://lnkd.in/emUePTiq
π₯10β€5π1πΎ1
This media is not supported in your browser
VIEW IN TELEGRAM
πͺ£Weather-Conditioned Depth Anythingπͺ£
πWeather-Conditioned Depth Anything from Texas A&M is the new SOTA in weather-robust depth estimation. A curated mix of real and synthetic degradation datasets to extract content-independent, degradation-aware weather embeddings. Repo under Apacheπ
πReview https://lnkd.in/p/eW-dsepD
πPaper https://lnkd.in/er_MvVft
πProject https://lnkd.in/ehXPs3C7
πRepo https://lnkd.in/edk7Ts_r
πWeather-Conditioned Depth Anything from Texas A&M is the new SOTA in weather-robust depth estimation. A curated mix of real and synthetic degradation datasets to extract content-independent, degradation-aware weather embeddings. Repo under Apacheπ
πReview https://lnkd.in/p/eW-dsepD
πPaper https://lnkd.in/er_MvVft
πProject https://lnkd.in/ehXPs3C7
πRepo https://lnkd.in/edk7Ts_r
π2π₯2β€1
π₯2π€―1
This media is not supported in your browser
VIEW IN TELEGRAM
π»Emerging Objs from Motionπ»
πMotion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0π
πReview https://lnkd.in/p/eezZrSJE
πPaper https://arxiv.org/pdf/2609.04348
πProject https://tj12342.github.io/object-concepts-from-motion/
πRepo https://github.com/TJ12342/object-concepts-from-motion/tree/main
πMotion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0π
πReview https://lnkd.in/p/eezZrSJE
πPaper https://arxiv.org/pdf/2609.04348
πProject https://tj12342.github.io/object-concepts-from-motion/
πRepo https://github.com/TJ12342/object-concepts-from-motion/tree/main
β€2π2π1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯#AIwithPapers: we are 17,000+π₯
π Even though 100+ bots are trying to join the discussion chats every day, there are 17,000 of us! Almost all of us are still humans π§
π Invite -> https://t.me/AI_DeepLearning
π Even though 100+ bots are trying to join the discussion chats every day, there are 17,000 of us! Almost all of us are still humans π§
π Invite -> https://t.me/AI_DeepLearning
β€24πΎ14π4
This media is not supported in your browser
VIEW IN TELEGRAM
πMcByte++ tracking-by-detectionπ
πMcByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0π
πReview https://lnkd.in/p/e4-diVJS
πPaper https://lnkd.in/e_Vxky-b
πRepo https://lnkd.in/e8SeCYmk
πMcByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0π
πReview https://lnkd.in/p/e4-diVJS
πPaper https://lnkd.in/e_Vxky-b
πRepo https://lnkd.in/e8SeCYmk
β€7π2π₯1π©1πΎ1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯π₯ Marigold V2 is out π₯π₯
πMarigold V2 is out: depth, (impressive) see-through depth, surface normals, albedo, and other dense modalities. SOTA results. Repo under Apache 2.0π
#AI #deeplearning #AIwithPapers
πReview https://lnkd.in/p/eKM44yDQ
πPaper https://arxiv.org/pdf/2609.08084
πRepo https://github.com/huawei-bayerlab/marigold-v2
πProject https://huggingface.co/spaces/huawei-bayerlab/marigold-v2-web
πMarigold V2 is out: depth, (impressive) see-through depth, surface normals, albedo, and other dense modalities. SOTA results. Repo under Apache 2.0π
#AI #deeplearning #AIwithPapers
πReview https://lnkd.in/p/eKM44yDQ
πPaper https://arxiv.org/pdf/2609.08084
πRepo https://github.com/huawei-bayerlab/marigold-v2
πProject https://huggingface.co/spaces/huawei-bayerlab/marigold-v2-web
π₯8π3β€2π1
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ΊEfficient/Scalable Video Pretrainingπ¦Ί
πLeVJEPA1 (Yann Lecun) is the first video encoder trained under LeJEPAβs collapse-free objective, and evaluate it under frozen probing against video and image pretraining baselines retrained on identical data, in both epoch-matched and FLOP-matched regimes. Repo under MITπ
πReview https://lnkd.in/p/eJQAm3AN
πPaper https://lnkd.in/eCzzTiNH
πProject https://levjepa.github.io/
πRepo https://lnkd.in/etiF5CDj
πLeVJEPA1 (Yann Lecun) is the first video encoder trained under LeJEPAβs collapse-free objective, and evaluate it under frozen probing against video and image pretraining baselines retrained on identical data, in both epoch-matched and FLOP-matched regimes. Repo under MITπ
πReview https://lnkd.in/p/eJQAm3AN
πPaper https://lnkd.in/eCzzTiNH
πProject https://levjepa.github.io/
πRepo https://lnkd.in/etiF5CDj
β€7π₯5π1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯ RelateAnything is gold! π₯
πRelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0π
πReview https://lnkd.in/p/etAcdFM3
πPaper https://arxiv.org/pdf/2609.12552
πRepo https://github.com/Maelic/RelateAnything
πProject https://maelic.github.io/RelateAnythingProject/
πRelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0π
πReview https://lnkd.in/p/etAcdFM3
πPaper https://arxiv.org/pdf/2609.12552
πRepo https://github.com/Maelic/RelateAnything
πProject https://maelic.github.io/RelateAnythingProject/
β€6π3π₯3π€―1