AI with Papers - Artificial Intelligence & Deep Learning
17K subscribers
161 photos
286 videos
14 files
1.48K links
All the AI with papers. Every day fresh updates about #DeepLearning #MachineLearning #LLM & #ComputerVision

Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/

#AI #chatGPT
Download Telegram
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ’„MagicMakeup TransferπŸ’„

πŸ‘‰Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Authors: Zhejiang University & vivo BlueImage Lab. Repo for non commercialπŸ’™

πŸ‘‰Review https://t.ly/JYpCr
πŸ‘‰Paper harxiv.org/pdf/2607.20924
πŸ‘‰Project vivocameraresearch.github.io/magicmakeup
πŸ‘‰Repo github.com/vivoCameraResearch/Magic-Makeup
❀4πŸ‘2πŸ”₯1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”ŽMicroZoom at Extreme ScaleπŸ”Ž

πŸ‘‰MicroZoom by UWA synthesizes gigapixel-resolution images grounded in consumer-grade microscope close-ups at magnification levels up to 350Γ—. Impressive. Repo under MITπŸ’™

πŸ‘‰Review https://t.ly/hgJD7
πŸ‘‰Paper https://arxiv.org/pdf/2607.24729
πŸ‘‰Project https://microzoom-sr.github.io/
πŸ‘‰Repo github.com/MicroZoom-SR/MicroZoom-SR.github.io
🀯5❀3πŸ‘2πŸ”₯1
This media is not supported in your browser
VIEW IN TELEGRAM
🍿 Dawn of Generative Cinematography 🍿

🟩 The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.

πŸ‘‰ Meanwhile, AI research is heading in the exact opposite direction.

🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.

πŸ‘‰More: https://t.ly/g-qUh
πŸ‘‰Paper arxiv.org/pdf/2607.24591
πŸ‘‰Project yixuanli98.github.io/cameraanything/
πŸ‘‰Repo github.com/yixuanli98/CameraAnything
πŸ”₯7❀5πŸ‘1πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯Decoder-only Any-to-Any ModelπŸ”₯

πŸ‘‰MODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under ApacheπŸ’™

πŸ‘‰Review https://t.ly/-2QKT
πŸ‘‰Paper https://lnkd.in/dhfBQGhB
πŸ‘‰Project https://lnkd.in/dPD_ECXk
πŸ‘‰Repo https://lnkd.in/dbDHw24u
πŸ”₯5❀3πŸ‘1🍾1
This media is not supported in your browser
VIEW IN TELEGRAM
🐠Dual-branch ID-Tracking🐠

πŸ‘‰TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT licenseπŸ’™

πŸ‘‰Review https://t.ly/WEDeY
πŸ‘‰Paper https://arxiv.org/pdf/2607.26412
πŸ‘‰Project https://vranlee.github.io/TIDE/
πŸ‘‰Repo https://github.com/vranlee/TIDE
πŸ”₯9❀7πŸ‘1πŸ‘1
πŸ”₯Unified Points n' LinesπŸ”₯

πŸ‘‰ETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under ApacheπŸ’™

πŸ‘‰Review https://lnkd.in/p/eW8j5JZj
πŸ‘‰Paper https://arxiv.org/pdf/2608.19894
πŸ‘‰Repo https://github.com/francois141/upal
❀11πŸ”₯7πŸ‘3🍾2πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ† Anyone in 4D is out πŸ†

πŸ‘‰4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/ec4dzGvb
πŸ‘‰Paper https://arxiv.org/pdf/2608.20335
πŸ‘‰Project https://4danyone.github.io
πŸ‘‰Repo github.com/ant-research/4DAnyone
πŸ”₯7❀4πŸ‘1πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‹β€πŸŸ©Remesh-Aware Mesh DeformationπŸ‹β€πŸŸ©

πŸ‘‰RADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MITπŸ’™

πŸ‘‰Review https://lnkd.in/p/eK6FZv9c
πŸ‘‰Paper https://arxiv.org/pdf/2608.17182
πŸ‘‰Project https://threedle.github.io/radmesh/
πŸ‘‰Repo https://github.com/threedle/radmesh/
❀5πŸ”₯3πŸ‘1πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ¦‘Unified Segmentation n' RetrievalπŸ¦‘

πŸ‘‰FoundYou gets an example of your object and it segments the same physical instance in a new image or retrieve it from a large gallery with ONE super-compact model. Repo/demo availableπŸ’™

πŸ‘‰Review https://lnkd.in/p/ex2qnKHW
πŸ‘‰Paper arxiv.org/pdf/2608.29917
πŸ‘‰Project https://lnkd.in/eNEUB_nV
πŸ‘‰Repo https://lnkd.in/eRyDY6Ue
πŸ”₯10❀4πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🍚Vision Weight Estimation🍚

πŸ‘‰Doppio is a novel video dataset capturing video of falling ground coffee, paired with precise, per-frame ground-truth weight measurements: OCR readings are extracted from the display, smoothed and time-lag compensated, and paired with per-frame weight annotations. Repo to be released under ApacheπŸ’™

πŸ‘‰Review https://www.linkedin.com/posts/visionarynet_computer-vision-weight-estimation-activity-7501900695457951744-ArBO
πŸ‘‰Paper https://lnkd.in/eHuy87SX
πŸ‘‰Project https://lnkd.in/e9g9zeK3
πŸ‘‰Repo https://lnkd.in/emUePTiq
πŸ”₯10❀5πŸ‘1🍾1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸͺ£Weather-Conditioned Depth AnythingπŸͺ£

πŸ‘‰Weather-Conditioned Depth Anything from Texas A&M is the new SOTA in weather-robust depth estimation. A curated mix of real and synthetic degradation datasets to extract content-independent, degradation-aware weather embeddings. Repo under ApacheπŸ’™

πŸ‘‰Review https://lnkd.in/p/eW-dsepD
πŸ‘‰Paper https://lnkd.in/er_MvVft
πŸ‘‰Project https://lnkd.in/ehXPs3C7
πŸ‘‰Repo https://lnkd.in/edk7Ts_r
πŸ‘2πŸ”₯2❀1
+++ Mistral raises 3B € +++

πŸ‘‰Discussion: https://lnkd.in/p/eVpF--VW
πŸ”₯2🀯1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‘»Emerging Objs from MotionπŸ‘»

πŸ‘‰Motion boundaries provide a strong signal for object-level grouping and can be used to derive pseudo-instance supervision. Suitable for: mono-depth, 3D object detection, 3D occupancy, and end-to-end planning. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/eezZrSJE
πŸ‘‰Paper https://arxiv.org/pdf/2609.04348
πŸ‘‰Project https://tj12342.github.io/object-concepts-from-motion/
πŸ‘‰Repo https://github.com/TJ12342/object-concepts-from-motion/tree/main
❀2πŸ‘2😍1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯#AIwithPapers: we are 17,000+πŸ”₯

πŸ‘‰ Even though 100+ bots are trying to join the discussion chats every day, there are 17,000 of us! Almost all of us are still humans 🧟

😈 Invite -> https://t.me/AI_DeepLearning
❀24🍾14πŸ‘4
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ€McByte++ tracking-by-detectionπŸ€

πŸ‘‰McByte++ is the newer extension of McByte that advances training-free sports MOT toward long-term ID tracking, while simultaneously improving efficiency and runtime performance. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/e4-diVJS
πŸ‘‰Paper https://lnkd.in/e_Vxky-b
πŸ‘‰Repo https://lnkd.in/e8SeCYmk
❀7πŸ‘2πŸ”₯1πŸ’©1🍾1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯πŸ”₯ Marigold V2 is out πŸ”₯πŸ”₯

πŸ‘‰Marigold V2 is out: depth, (impressive) see-through depth, surface normals, albedo, and other dense modalities. SOTA results. Repo under Apache 2.0πŸ’™

#AI #deeplearning #AIwithPapers

πŸ‘‰Review https://lnkd.in/p/eKM44yDQ
πŸ‘‰Paper https://arxiv.org/pdf/2609.08084
πŸ‘‰Repo https://github.com/huawei-bayerlab/marigold-v2
πŸ‘‰Project https://huggingface.co/spaces/huawei-bayerlab/marigold-v2-web
πŸ”₯8πŸ‘3❀2πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🦺Efficient/Scalable Video Pretraining🦺

πŸ‘‰LeVJEPA1 (Yann Lecun) is the first video encoder trained under LeJEPA’s collapse-free objective, and evaluate it under frozen probing against video and image pretraining baselines retrained on identical data, in both epoch-matched and FLOP-matched regimes. Repo under MITπŸ’™

πŸ‘‰Review https://lnkd.in/p/eJQAm3AN
πŸ‘‰Paper https://lnkd.in/eCzzTiNH
πŸ‘‰Project https://levjepa.github.io/
πŸ‘‰Repo https://lnkd.in/etiF5CDj
❀7πŸ”₯5πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯ RelateAnything is gold! πŸ”₯

πŸ‘‰RelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/etAcdFM3
πŸ‘‰Paper https://arxiv.org/pdf/2609.12552
πŸ‘‰Repo https://github.com/Maelic/RelateAnything
πŸ‘‰Project https://maelic.github.io/RelateAnythingProject/
❀6πŸ‘3πŸ”₯3🀯1