This media is not supported in your browser
VIEW IN TELEGRAM
π¦Unified Segmentation n' Retrievalπ¦
πFoundYou gets an example of your object and it segments the same physical instance in a new image or retrieve it from a large gallery with ONE super-compact model. Repo/demo availableπ
πReview https://lnkd.in/p/ex2qnKHW
πPaper arxiv.org/pdf/2608.29917
πProject https://lnkd.in/eNEUB_nV
πRepo https://lnkd.in/eRyDY6Ue
πFoundYou gets an example of your object and it segments the same physical instance in a new image or retrieve it from a large gallery with ONE super-compact model. Repo/demo availableπ
πReview https://lnkd.in/p/ex2qnKHW
πPaper arxiv.org/pdf/2608.29917
πProject https://lnkd.in/eNEUB_nV
πRepo https://lnkd.in/eRyDY6Ue
π₯10β€4π1
This media is not supported in your browser
VIEW IN TELEGRAM
πVision Weight Estimationπ
πDoppio is a novel video dataset capturing video of falling ground coffee, paired with precise, per-frame ground-truth weight measurements: OCR readings are extracted from the display, smoothed and time-lag compensated, and paired with per-frame weight annotations. Repo to be released under Apacheπ
πReview https://www.linkedin.com/posts/visionarynet_computer-vision-weight-estimation-activity-7501900695457951744-ArBO
πPaper https://lnkd.in/eHuy87SX
πProject https://lnkd.in/e9g9zeK3
πRepo https://lnkd.in/emUePTiq
πDoppio is a novel video dataset capturing video of falling ground coffee, paired with precise, per-frame ground-truth weight measurements: OCR readings are extracted from the display, smoothed and time-lag compensated, and paired with per-frame weight annotations. Repo to be released under Apacheπ
πReview https://www.linkedin.com/posts/visionarynet_computer-vision-weight-estimation-activity-7501900695457951744-ArBO
πPaper https://lnkd.in/eHuy87SX
πProject https://lnkd.in/e9g9zeK3
πRepo https://lnkd.in/emUePTiq
π₯10β€5π1πΎ1
This media is not supported in your browser
VIEW IN TELEGRAM
πͺ£Weather-Conditioned Depth Anythingπͺ£
πWeather-Conditioned Depth Anything from Texas A&M is the new SOTA in weather-robust depth estimation. A curated mix of real and synthetic degradation datasets to extract content-independent, degradation-aware weather embeddings. Repo under Apacheπ
πReview https://lnkd.in/p/eW-dsepD
πPaper https://lnkd.in/er_MvVft
πProject https://lnkd.in/ehXPs3C7
πRepo https://lnkd.in/edk7Ts_r
πWeather-Conditioned Depth Anything from Texas A&M is the new SOTA in weather-robust depth estimation. A curated mix of real and synthetic degradation datasets to extract content-independent, degradation-aware weather embeddings. Repo under Apacheπ
πReview https://lnkd.in/p/eW-dsepD
πPaper https://lnkd.in/er_MvVft
πProject https://lnkd.in/ehXPs3C7
πRepo https://lnkd.in/edk7Ts_r
π2π₯2β€1
π₯2π€―1
This media is not supported in your browser
VIEW IN TELEGRAM
π»Emerging Objs from Motionπ»
πMotion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0π
πReview https://lnkd.in/p/eezZrSJE
πPaper https://arxiv.org/pdf/2609.04348
πProject https://tj12342.github.io/object-concepts-from-motion/
πRepo https://github.com/TJ12342/object-concepts-from-motion/tree/main
πMotion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0π
πReview https://lnkd.in/p/eezZrSJE
πPaper https://arxiv.org/pdf/2609.04348
πProject https://tj12342.github.io/object-concepts-from-motion/
πRepo https://github.com/TJ12342/object-concepts-from-motion/tree/main
β€2π2π1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯#AIwithPapers: we are 17,000+π₯
π Even though 100+ bots are trying to join the discussion chats every day, there are 17,000 of us! Almost all of us are still humans π§
π Invite -> https://t.me/AI_DeepLearning
π Even though 100+ bots are trying to join the discussion chats every day, there are 17,000 of us! Almost all of us are still humans π§
π Invite -> https://t.me/AI_DeepLearning
β€25πΎ14π4
This media is not supported in your browser
VIEW IN TELEGRAM
πMcByte++ tracking-by-detectionπ
πMcByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0π
πReview https://lnkd.in/p/e4-diVJS
πPaper https://lnkd.in/e_Vxky-b
πRepo https://lnkd.in/e8SeCYmk
πMcByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0π
πReview https://lnkd.in/p/e4-diVJS
πPaper https://lnkd.in/e_Vxky-b
πRepo https://lnkd.in/e8SeCYmk
β€8π2π₯1π©1πΎ1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯π₯ Marigold V2 is out π₯π₯
πMarigold V2 is out: depth, (impressive) see-through depth, surface normals, albedo, and other dense modalities. SOTA results. Repo under Apache 2.0π
#AI #deeplearning #AIwithPapers
πReview https://lnkd.in/p/eKM44yDQ
πPaper https://arxiv.org/pdf/2609.08084
πRepo https://github.com/huawei-bayerlab/marigold-v2
πProject https://huggingface.co/spaces/huawei-bayerlab/marigold-v2-web
πMarigold V2 is out: depth, (impressive) see-through depth, surface normals, albedo, and other dense modalities. SOTA results. Repo under Apache 2.0π
#AI #deeplearning #AIwithPapers
πReview https://lnkd.in/p/eKM44yDQ
πPaper https://arxiv.org/pdf/2609.08084
πRepo https://github.com/huawei-bayerlab/marigold-v2
πProject https://huggingface.co/spaces/huawei-bayerlab/marigold-v2-web
π₯9β€3π3π1
This media is not supported in your browser
VIEW IN TELEGRAM
π¦ΊEfficient/Scalable Video Pretrainingπ¦Ί
πLeVJEPA1 (Yann Lecun) is the first video encoder trained under LeJEPAβs collapse-free objective, and evaluate it under frozen probing against video and image pretraining baselines retrained on identical data, in both epoch-matched and FLOP-matched regimes. Repo under MITπ
πReview https://lnkd.in/p/eJQAm3AN
πPaper https://lnkd.in/eCzzTiNH
πProject https://levjepa.github.io/
πRepo https://lnkd.in/etiF5CDj
πLeVJEPA1 (Yann Lecun) is the first video encoder trained under LeJEPAβs collapse-free objective, and evaluate it under frozen probing against video and image pretraining baselines retrained on identical data, in both epoch-matched and FLOP-matched regimes. Repo under MITπ
πReview https://lnkd.in/p/eJQAm3AN
πPaper https://lnkd.in/eCzzTiNH
πProject https://levjepa.github.io/
πRepo https://lnkd.in/etiF5CDj
β€8π₯5π1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯ RelateAnything is gold! π₯
πRelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0π
πReview https://lnkd.in/p/etAcdFM3
πPaper https://arxiv.org/pdf/2609.12552
πRepo https://github.com/Maelic/RelateAnything
πProject https://maelic.github.io/RelateAnythingProject/
πRelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0π
πReview https://lnkd.in/p/etAcdFM3
πPaper https://arxiv.org/pdf/2609.12552
πRepo https://github.com/Maelic/RelateAnything
πProject https://maelic.github.io/RelateAnythingProject/
β€12π₯4π3π1π€―1
This media is not supported in your browser
VIEW IN TELEGRAM
π EventEgoHands++ is out! π
πEventEgoHands++ is a novel framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint. 1M+ samples dataset! Code/Data releasedπ
πReview https://lnkd.in/p/eTbPvXbW
πPaper https://arxiv.org/pdf/2609.17189
πRepo https://github.com/ryhara/EventEgoHandsV2
πProject https://ryhara.github.io/EventEgoHandsV2/
πEventEgoHands++ is a novel framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint. 1M+ samples dataset! Code/Data releasedπ
πReview https://lnkd.in/p/eTbPvXbW
πPaper https://arxiv.org/pdf/2609.17189
πRepo https://github.com/ryhara/EventEgoHandsV2
πProject https://ryhara.github.io/EventEgoHandsV2/
β€3π₯2π1
This media is not supported in your browser
VIEW IN TELEGRAM
π¦SOTA Splashing Liquidsπ¦
πSplashSplat reconstructs splashing liquids from real multi-view vide. Impose physical structure only where the observations can constrain it. Impressive results, SOTA. Code TBR under MITπ
πReview https://lnkd.in/p/ejMTHcp7
πPaper https://arxiv.org/pdf/2609.20818
πProject niko-creater.github.io/splashsplat-web/
πRepo https://github.com/Niko-creater/Splashsplat
πSplashSplat reconstructs splashing liquids from real multi-view vide. Impose physical structure only where the observations can constrain it. Impressive results, SOTA. Code TBR under MITπ
πReview https://lnkd.in/p/ejMTHcp7
πPaper https://arxiv.org/pdf/2609.20818
πProject niko-creater.github.io/splashsplat-web/
πRepo https://github.com/Niko-creater/Splashsplat
π4β€2π₯2π1
This media is not supported in your browser
VIEW IN TELEGRAM
π₯Agentic Image-to-Sceneπ₯
πHARMONY by UPenn is a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Impressive 3D scenes. Repo TBAπ
πReview https://lnkd.in/p/ep2hmRSp
πPaper https://arxiv.org/pdf/2609.26793
πProject https://cwchenwang.github.io/harmony/
πData https://huggingface.co/datasets/ShufanSun/harmony
πHARMONY by UPenn is a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Impressive 3D scenes. Repo TBAπ
πReview https://lnkd.in/p/ep2hmRSp
πPaper https://arxiv.org/pdf/2609.26793
πProject https://cwchenwang.github.io/harmony/
πData https://huggingface.co/datasets/ShufanSun/harmony
π₯8β€2π1
This media is not supported in your browser
VIEW IN TELEGRAM
πΏPanoSeg3R: SOTA 3D SegmentationπΏ
πPanoSeg3R is a novel feed-forward framework for 3D panoramic semantic segmentation. New SOTA. Code comingπ
πReview https://lnkd.in/p/eKCKWv3g
πPaper https://arxiv.org/pdf/2609.22687
πProject https://harryyoon777.github.io/PanoSeg3R/#
πRepo TBA
πPanoSeg3R is a novel feed-forward framework for 3D panoramic semantic segmentation. New SOTA. Code comingπ
πReview https://lnkd.in/p/eKCKWv3g
πPaper https://arxiv.org/pdf/2609.22687
πProject https://harryyoon777.github.io/PanoSeg3R/#
πRepo TBA
β€4π1π₯1π1
This media is not supported in your browser
VIEW IN TELEGRAM
π©»Universal X-ray Segmentationπ©»
πFleXray: universal anatomical segmentation across the entire body in clinical X-rays. Built on a scalable, physics-based generative X-ray data engine. Repo under MITπ
πReview https://lnkd.in/p/e9MUk_eq
πPaper https://arxiv.org/pdf/2609.26756
πProject https://flexray.csail.mit.edu/
πRepo https://github.com/VictorButoi/FleXray
πFleXray: universal anatomical segmentation across the entire body in clinical X-rays. Built on a scalable, physics-based generative X-ray data engine. Repo under MITπ
πReview https://lnkd.in/p/e9MUk_eq
πPaper https://arxiv.org/pdf/2609.26756
πProject https://flexray.csail.mit.edu/
πRepo https://github.com/VictorButoi/FleXray
π4β€3π₯2π2
This media is not supported in your browser
VIEW IN TELEGRAM
π¦΄3D Foundational Radiologyπ¦΄
πnnFoundation: 3D radiological foundation models designed for transferable representation learning across heterogeneous tasks/datasets. Models releasedπ
πReview https://lnkd.in/p/eNajRGBi
πPaper https://arxiv.org/pdf/2609.26924
πModels https://huggingface.co/collections/MIC-DKFZ/nnfoundation
πnnFoundation: 3D radiological foundation models designed for transferable representation learning across heterogeneous tasks/datasets. Models releasedπ
πReview https://lnkd.in/p/eNajRGBi
πPaper https://arxiv.org/pdf/2609.26924
πModels https://huggingface.co/collections/MIC-DKFZ/nnfoundation
β€8π2π2π€©1