DS & ML
589 subscribers
1.17K photos
238 videos
240 files
1.41K links
Download Telegram
​🐍 Распознавание текста с изображения — [8:50]

Python часто используют при разработке искусственного интеллекта, который, в свою очередь, умеет в компьютерное зрение. Это самое зрение позволяет программе искать и идентифицировать объекты на изображении (люди, животные или даже текст), без чего, естественно, никак не обойтись.

В этом видео автор на практике показал, как с использованием EasyOCR считывать текст русского и английского языка с изображения, а после — записать его в файл.

Перейти к просмотру

#видео #python
Forwarded from Big Data Science
🔥Hide a spicy photo from strangers? Easy with nudity detection API
Python program by DeepAI
(https://deepai.org/machine-learning-model/nsfw-detector)
evaluates the image and estimates the likelihood that it covers areas of the human body that are usually found under clothing. The nudity check is a dynamic label, a certainty algorithm, or a false percentage. The user can set a different threshold in their app for what it considers to be nudity, and the image detection algorithm will return the percentage chance that the image contains natural content. This ML system is applicable not only to photos, but also to videos: neural networks analyze the video stream and give probabilistic feedback on the “adultness” of consumption.
See an example of using the API here: https://medium.com/mlearning-ai/i-tried-a-python-nude-detector-with-my-photo-446dba1bbfc8
Forwarded from GitHub Community
​TextSnatcher – Инструмент, что позволяет копировать текст из изображений в буфер обмена за считанные секунды

К сожалению работает только на Linux

GitHub | #Linux #OCR #Useful
DS & ML pinned «​TextSnatcher – Инструмент, что позволяет копировать текст из изображений в буфер обмена за считанные секунды К сожалению работает только на Linux GitHub | #Linux #OCR #Useful»
Forwarded from тоже моушн
This media is not supported in your browser
VIEW IN TELEGRAM
с праздником, пацаны)

в комментариях к посту про музыкальный StyleGAN3 спрашивали - можно ли такое провернуть со своими фотографиями? теперь можно! для этого надо сначала заморфить фотки в колабе frame interpolation, а затем полученное видео обработать под звук в новом колабе - audio reactive video

make love not war!

colab audio reactive video
colab sequence frame interpolation
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Интересное от Google Research

"Известно, что StyleGAN создает высококачественные изображения, а также предлагает беспрецедентное семантическое редактирование. Однако эти удивительные возможности были продемонстрированы только на ограниченном наборе наборов данных, которые обычно структурно выровнены и тщательно отобраны. В этой статье мы покажем, как StyleGAN можно адаптировать для работы с необработанными изображениями, собранными из Интернета."

Self-Distilled StyleGAN: Towards Generation from Internet Photos https://self-distilled-stylegan.github.io/
github:
https://github.com/self-distilled-stylegan/self-distilled-internet-photos
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Примеры семантического редактирования:

"Наш подход сохраняет замечательные возможности семантического редактирования StyleGAN, что позволяет нам выполнять различное редактирование для каждого домена."
Forwarded from Big Data Science
😎3 Face Recognition ML Services APIs: Choose What You Need
• IBM Watson Visual Recognition API for identifying scenes, objects, and faces in images uploaded to the service. It can process unstructured data in a large volume and is suitable as a decision support system. But it is expensive to maintain and does not process structured data directly. The facial recognition method does not support general biometric recognition, and the maximum image size is 10 MB with a minimum recommended density of 32x32 ppi. Suitable for image classification using built-in classifiers, allows you to create your own classifiers and train ML models. https://www.ibm.com/watson
• Kairos Face Recognition API allows developers of ML applications to add face recognition capabilities to their applications by writing just a few lines of code. The Kairos Face Recognition API shows high accuracy in real-life scenarios and performs well in low light conditions as well as partial face hiding. Applies an ethical approach to identifying individuals, taking into account diversity. It is an extensible tool: users can apply additional intelligence to work with video and photos in the real world. Suitable for working with large volumes of images and ensures confidentiality through the secure storage of collected data and regular audits. However, it only supports BMP, JPG, and PNG file types, GIF files are not supported. Slightly slower in operation than the AWS API. https://www.kairos.com/docs/getting-started-with-kairos-face-recognition
• Microsoft Computer Vision API in Azure gives developers access to advanced image processing algorithms. Once an image is loaded or its URL is specified, Microsoft Computer Vision algorithms analyze its visual content in various ways based on the user's choice. An added benefit of this fast API is visual guides, tutorials, and examples. A high SLA guarantees at least 99.9% availability. Through tight integration with other Microsoft Azure cloud services, APIs can be packaged into a complete solution. But if the transaction per second limit is exceeded, the response time will be reduced to the agreed limit. The pricing model is demand-driven, so the service can become expensive if the number of requests spikes. The Microsoft Computer Vision API is great for classifying images with objects, creatures, scenery, and activities, including their identification, categorization, and image tagging. Supports face, mood, age and scene recognition, optical character recognition to detect text content in images. Also provides intelligent photo management and moderated content display restriction. https://azure.microsoft.com/en-us/services/cognitive-services/computer-vision/
Forwarded from GitHub Community
​Larynx – Автономная сквозная система преобразования текста в речь с использованием gruut и onnx (архитектура). Доступно 50 голосов на 9 языках.

• Достаточно хороший синтез, чтобы не использовать облачный сервис
• Быстрее, чем работа в реальном времени на Raspberry Pi 4 (с вокодером низкого качества)
• Широкая языковая поддержка (9 языков)
• Голоса обучены исключительно на основе публичных наборов данных

GitHub | #Python #Speech #Interesting
Forwarded from [PYTHON:TODAY]
🔥 Полезные библиотеки Python

DeepFaceLive
- Python утилита для создания дип фейков в режиме реального времени.

⚙️ GitHub/Инструкция

💾 Больше интересных проектов

#python #github
Forwarded from Derp Learning
This media is not supported in your browser
VIEW IN TELEGRAM
DiscoDiffusion v5 3d Работает неплохо даже на дефолтных настройках.

Забавны последние кадры, где все варианты CLIP решили, что голубой фон достаточно похож на текстовый запрос, и рисовать мы ничего больше не будем.
Forwarded from эйай ньюз
This media is not supported in your browser
VIEW IN TELEGRAM
Google прокачал StyleGAN

Теперь, наконец, нормально генерятся не только лица. Обучают на неструктурированных картинках из интернета, и результат вы можете видеть на видео.

https://self-distilled-stylegan.github.io/
Forwarded from GitHub Community
​Keras – это API для глубокого обучения, написанный на Python, и работающий поверх платформы машинного обучения TensorFlow

Он был разработан с акцентом на быстродействие. Возможность перейти от идеи к результату как можно быстрее является ключом к проведению хороших исследований.

GitHub | #Python #Deep #Learning #API
Forwarded from GitHub Community
​DeepFaceLive – Python утилита для создания дипфейков в режиме реального времени

Минимальные системные требования:
• Любая видеокарта, совместимая с DirectX12
• Современный процессор с инструкциями AVX
• 4 ГБ оперативной памяти, файл подкачки 32 ГБ+

GitHub | #Python #Deep #Fake #Interesting
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Cryptopunks GAN

These CryptoPunks do not exist. 👀

Изображение было создано с помощью DCGAN, обученной на изображениях криптопанков.
https://github.com/teddykoker/cryptopunks-gan

Попробовать: https://huggingface.co/spaces/nateraw/cryptopunks-generator
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Вышла новая версия DiscoDiffusion v5 3d
• Автор генерации
• Colab
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Интересное применение deepfacelive.

Окно просмотра maya транслируется через OBS и выводится в DeepFaceLive.

*Maya — программное обеспечение для 3D-анимации с мощными инструментами моделирования.