AI with Papers - Artificial Intelligence & Deep Learning
17K subscribers
161 photos
284 videos
14 files
1.47K links
All the AI with papers. Every day fresh updates about #DeepLearning #MachineLearning #LLM & #ComputerVision

Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/

#AI #chatGPT
Download Telegram
This media is not supported in your browser
VIEW IN TELEGRAM
🦜Streaming 4D Transformer🦜

👉IGGT4D is a novel a streaming instance-grounded geometry transformer for online 4D scene understanding. It processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. Repo/Data announced💙

👉Review https://t.ly/LFrKR
👉Paper https://arxiv.org/pdf/2607.19228
👉Project https://iggt4d.github.io/
👉Repo TBA
3🔥2👏1
🫛Spatially-Aware Class-Agnostic Counting🫛

👉UpCount is reference-free, spatially aware, class-agnostic object counting with an MAE-pretrained ViT, DPT-style feature reassembly, FeatUp-style joint bilateral upsampling & proposal verification. Repo MIT💙

👉Review https://t.ly/dWOc3
👉Paper https://arxiv.org/pdf/2607.16826
👉Repo github.com/r28112072-rgb/upcount
🔥52👍1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
💢Unified Video Dense Prediction💢

👉UniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBR💙

👉Review https://t.ly/oo7et
👉Paper https://arxiv.org/pdf/2607.21592
👉Project https://unid-video.github.io/
👉Repo https://github.com/YihongSun/UniD
🔥75👏1
This media is not supported in your browser
VIEW IN TELEGRAM
💄MagicMakeup Transfer💄

👉Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Authors: Zhejiang University & vivo BlueImage Lab. Repo for non commercial💙

👉Review https://t.ly/JYpCr
👉Paper harxiv.org/pdf/2607.20924
👉Project vivocameraresearch.github.io/magicmakeup
👉Repo github.com/vivoCameraResearch/Magic-Makeup
3👍1🔥1
This media is not supported in your browser
VIEW IN TELEGRAM
🔎MicroZoom at Extreme Scale🔎

👉MicroZoom by UWA synthesizes gigapixel-resolution images grounded in consumer-grade microscope close-ups at magnification levels up to 350×. Impressive. Repo under MIT💙

👉Review https://t.ly/hgJD7
👉Paper https://arxiv.org/pdf/2607.24729
👉Project https://microzoom-sr.github.io/
👉Repo github.com/MicroZoom-SR/MicroZoom-SR.github.io
🤯53🔥1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
🍿 Dawn of Generative Cinematography 🍿

🟩 The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture.

👉 Meanwhile, AI research is heading in the exact opposite direction.

🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere.

👉More: https://t.ly/g-qUh
👉Paper arxiv.org/pdf/2607.24591
👉Project yixuanli98.github.io/cameraanything/
👉Repo github.com/yixuanli98/CameraAnything
🔥75👍1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
🔥Decoder-only Any-to-Any Model🔥

👉MODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under Apache💙

👉Review https://t.ly/-2QKT
👉Paper https://lnkd.in/dhfBQGhB
👉Project https://lnkd.in/dPD_ECXk
👉Repo https://lnkd.in/dbDHw24u
🔥53👍1🍾1
This media is not supported in your browser
VIEW IN TELEGRAM
🐠Dual-branch ID-Tracking🐠

👉TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT license💙

👉Review https://t.ly/WEDeY
👉Paper https://arxiv.org/pdf/2607.26412
👉Project https://vranlee.github.io/TIDE/
👉Repo https://github.com/vranlee/TIDE
🔥96👍1👏1
🔥Unified Points n' Lines🔥

👉ETH (+Microsoft Spatial AI Lab) unveils a novel feature extractor that jointly extracts keypoints, lines, and feature descriptors within a single lightweight net. SOTA in line detection can be achieved by adding only three convolutional layers to existing point extractor. Repo under Apache💙

👉Review https://lnkd.in/p/eW8j5JZj
👉Paper https://arxiv.org/pdf/2608.19894
👉Repo https://github.com/francois141/upal
10🔥6👍3👏1🍾1
This media is not supported in your browser
VIEW IN TELEGRAM
🐆 Anyone in 4D is out 🐆

👉4DAnyone turns a casual monocular video into multi-view videos, enabling downstream 4DGS reconstruction. Full repo under Apache 2.0💙

👉Review https://lnkd.in/p/ec4dzGvb
👉Paper https://arxiv.org/pdf/2608.20335
👉Project https://4danyone.github.io
👉Repo github.com/ant-research/4DAnyone
🔥72👍1👏1
This media is not supported in your browser
VIEW IN TELEGRAM
🍋‍🟩Remesh-Aware Mesh Deformation🍋‍🟩

👉RADmesh is a novel generative deformation technique enhanced by remeshing. Given a text prompt, it deforms and remeshes a mesh region to form new geometric features. Repo MIT💙

👉Review https://lnkd.in/p/eK6FZv9c
👉Paper https://arxiv.org/pdf/2608.17182
👉Project https://threedle.github.io/radmesh/
👉Repo https://github.com/threedle/radmesh/
🔥32👍1👏1