Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Интересное от Google Research
"Известно, что StyleGAN создает высококачественные изображения, а также предлагает беспрецедентное семантическое редактирование. Однако эти удивительные возможности были продемонстрированы только на ограниченном наборе наборов данных, которые обычно структурно выровнены и тщательно отобраны. В этой статье мы покажем, как StyleGAN можно адаптировать для работы с необработанными изображениями, собранными из Интернета."
Self-Distilled StyleGAN: Towards Generation from Internet Photos https://self-distilled-stylegan.github.io/
github:
https://github.com/self-distilled-stylegan/self-distilled-internet-photos
"Известно, что StyleGAN создает высококачественные изображения, а также предлагает беспрецедентное семантическое редактирование. Однако эти удивительные возможности были продемонстрированы только на ограниченном наборе наборов данных, которые обычно структурно выровнены и тщательно отобраны. В этой статье мы покажем, как StyleGAN можно адаптировать для работы с необработанными изображениями, собранными из Интернета."
Self-Distilled StyleGAN: Towards Generation from Internet Photos https://self-distilled-stylegan.github.io/
github:
https://github.com/self-distilled-stylegan/self-distilled-internet-photos
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Примеры семантического редактирования:
"Наш подход сохраняет замечательные возможности семантического редактирования StyleGAN, что позволяет нам выполнять различное редактирование для каждого домена."
"Наш подход сохраняет замечательные возможности семантического редактирования StyleGAN, что позволяет нам выполнять различное редактирование для каждого домена."
Forwarded from Big Data Science
😎3 Face Recognition ML Services APIs: Choose What You Need
• IBM Watson Visual Recognition API for identifying scenes, objects, and faces in images uploaded to the service. It can process unstructured data in a large volume and is suitable as a decision support system. But it is expensive to maintain and does not process structured data directly. The facial recognition method does not support general biometric recognition, and the maximum image size is 10 MB with a minimum recommended density of 32x32 ppi. Suitable for image classification using built-in classifiers, allows you to create your own classifiers and train ML models. https://www.ibm.com/watson
• Kairos Face Recognition API allows developers of ML applications to add face recognition capabilities to their applications by writing just a few lines of code. The Kairos Face Recognition API shows high accuracy in real-life scenarios and performs well in low light conditions as well as partial face hiding. Applies an ethical approach to identifying individuals, taking into account diversity. It is an extensible tool: users can apply additional intelligence to work with video and photos in the real world. Suitable for working with large volumes of images and ensures confidentiality through the secure storage of collected data and regular audits. However, it only supports BMP, JPG, and PNG file types, GIF files are not supported. Slightly slower in operation than the AWS API. https://www.kairos.com/docs/getting-started-with-kairos-face-recognition
• Microsoft Computer Vision API in Azure gives developers access to advanced image processing algorithms. Once an image is loaded or its URL is specified, Microsoft Computer Vision algorithms analyze its visual content in various ways based on the user's choice. An added benefit of this fast API is visual guides, tutorials, and examples. A high SLA guarantees at least 99.9% availability. Through tight integration with other Microsoft Azure cloud services, APIs can be packaged into a complete solution. But if the transaction per second limit is exceeded, the response time will be reduced to the agreed limit. The pricing model is demand-driven, so the service can become expensive if the number of requests spikes. The Microsoft Computer Vision API is great for classifying images with objects, creatures, scenery, and activities, including their identification, categorization, and image tagging. Supports face, mood, age and scene recognition, optical character recognition to detect text content in images. Also provides intelligent photo management and moderated content display restriction. https://azure.microsoft.com/en-us/services/cognitive-services/computer-vision/
• IBM Watson Visual Recognition API for identifying scenes, objects, and faces in images uploaded to the service. It can process unstructured data in a large volume and is suitable as a decision support system. But it is expensive to maintain and does not process structured data directly. The facial recognition method does not support general biometric recognition, and the maximum image size is 10 MB with a minimum recommended density of 32x32 ppi. Suitable for image classification using built-in classifiers, allows you to create your own classifiers and train ML models. https://www.ibm.com/watson
• Kairos Face Recognition API allows developers of ML applications to add face recognition capabilities to their applications by writing just a few lines of code. The Kairos Face Recognition API shows high accuracy in real-life scenarios and performs well in low light conditions as well as partial face hiding. Applies an ethical approach to identifying individuals, taking into account diversity. It is an extensible tool: users can apply additional intelligence to work with video and photos in the real world. Suitable for working with large volumes of images and ensures confidentiality through the secure storage of collected data and regular audits. However, it only supports BMP, JPG, and PNG file types, GIF files are not supported. Slightly slower in operation than the AWS API. https://www.kairos.com/docs/getting-started-with-kairos-face-recognition
• Microsoft Computer Vision API in Azure gives developers access to advanced image processing algorithms. Once an image is loaded or its URL is specified, Microsoft Computer Vision algorithms analyze its visual content in various ways based on the user's choice. An added benefit of this fast API is visual guides, tutorials, and examples. A high SLA guarantees at least 99.9% availability. Through tight integration with other Microsoft Azure cloud services, APIs can be packaged into a complete solution. But if the transaction per second limit is exceeded, the response time will be reduced to the agreed limit. The pricing model is demand-driven, so the service can become expensive if the number of requests spikes. The Microsoft Computer Vision API is great for classifying images with objects, creatures, scenery, and activities, including their identification, categorization, and image tagging. Supports face, mood, age and scene recognition, optical character recognition to detect text content in images. Also provides intelligent photo management and moderated content display restriction. https://azure.microsoft.com/en-us/services/cognitive-services/computer-vision/
Ibm
IBM Watson
See how IBM Watson has advanced enterprise AI.
Forwarded from GitHub Community
Larynx – Автономная сквозная система преобразования текста в речь с использованием gruut и onnx (архитектура). Доступно 50 голосов на 9 языках.
• Достаточно хороший синтез, чтобы не использовать облачный сервис
• Быстрее, чем работа в реальном времени на Raspberry Pi 4 (с вокодером низкого качества)
• Широкая языковая поддержка (9 языков)
• Голоса обучены исключительно на основе публичных наборов данных
GitHub | #Python #Speech #Interesting
• Достаточно хороший синтез, чтобы не использовать облачный сервис
• Быстрее, чем работа в реальном времени на Raspberry Pi 4 (с вокодером низкого качества)
• Широкая языковая поддержка (9 языков)
• Голоса обучены исключительно на основе публичных наборов данных
GitHub | #Python #Speech #Interesting
Forwarded from [PYTHON:TODAY]
🔥 Полезные библиотеки Python
DeepFaceLive - Python утилита для создания дип фейков в режиме реального времени.
⚙️ GitHub/Инструкция
💾 Больше интересных проектов
#python #github
DeepFaceLive - Python утилита для создания дип фейков в режиме реального времени.
⚙️ GitHub/Инструкция
💾 Больше интересных проектов
#python #github
Forwarded from Derp Learning
This media is not supported in your browser
VIEW IN TELEGRAM
DiscoDiffusion v5 3d Работает неплохо даже на дефолтных настройках.
Забавны последние кадры, где все варианты CLIP решили, что голубой фон достаточно похож на текстовый запрос, и рисовать мы ничего больше не будем.
Забавны последние кадры, где все варианты CLIP решили, что голубой фон достаточно похож на текстовый запрос, и рисовать мы ничего больше не будем.
Forwarded from эйай ньюз
This media is not supported in your browser
VIEW IN TELEGRAM
Google прокачал StyleGAN
Теперь, наконец, нормально генерятся не только лица. Обучают на неструктурированных картинках из интернета, и результат вы можете видеть на видео.
https://self-distilled-stylegan.github.io/
Теперь, наконец, нормально генерятся не только лица. Обучают на неструктурированных картинках из интернета, и результат вы можете видеть на видео.
https://self-distilled-stylegan.github.io/
Forwarded from GitHub Community
Keras – это API для глубокого обучения, написанный на Python, и работающий поверх платформы машинного обучения TensorFlow
Он был разработан с акцентом на быстродействие. Возможность перейти от идеи к результату как можно быстрее является ключом к проведению хороших исследований.
GitHub | #Python #Deep #Learning #API
Он был разработан с акцентом на быстродействие. Возможность перейти от идеи к результату как можно быстрее является ключом к проведению хороших исследований.
GitHub | #Python #Deep #Learning #API
Forwarded from GitHub Community
DeepFaceLive – Python утилита для создания дипфейков в режиме реального времени
Минимальные системные требования:
• Любая видеокарта, совместимая с DirectX12
• Современный процессор с инструкциями AVX
• 4 ГБ оперативной памяти, файл подкачки 32 ГБ+
GitHub | #Python #Deep #Fake #Interesting
Минимальные системные требования:
• Любая видеокарта, совместимая с DirectX12
• Современный процессор с инструкциями AVX
• 4 ГБ оперативной памяти, файл подкачки 32 ГБ+
GitHub | #Python #Deep #Fake #Interesting
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Cryptopunks GAN
These CryptoPunks do not exist. 👀
Изображение было создано с помощью DCGAN, обученной на изображениях криптопанков.
https://github.com/teddykoker/cryptopunks-gan
Попробовать: https://huggingface.co/spaces/nateraw/cryptopunks-generator
These CryptoPunks do not exist. 👀
Изображение было создано с помощью DCGAN, обученной на изображениях криптопанков.
https://github.com/teddykoker/cryptopunks-gan
Попробовать: https://huggingface.co/spaces/nateraw/cryptopunks-generator
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Forwarded from Технологии | Нейросети | Боты
This media is not supported in your browser
VIEW IN TELEGRAM
Интересное применение deepfacelive.
Окно просмотра maya транслируется через OBS и выводится в DeepFaceLive.
*Maya — программное обеспечение для 3D-анимации с мощными инструментами моделирования.
Окно просмотра maya транслируется через OBS и выводится в DeepFaceLive.
*Maya — программное обеспечение для 3D-анимации с мощными инструментами моделирования.
Forwarded from Big Data Science
👍🏻Sentiment analysis in social networks in Python with VADER without developing an ML model
Not every classification problem needs machine learning models: sometimes even simple approaches can give excellent results. For example, VADER (Valence Aware Dictionary and sEntiment Reasoner) is a vocabulary and rule based sentiment analysis model. The project source code is available on Github under the MIT license: https://github.com/cjhutto/vaderSentiment
VADER can efficiently handle dictionaries, abbreviations, capital letters, repetitive punctuation marks, emoticons (😢 , 😃 , 😭, etc.), etc., which are commonly used in social networks to express sentiment, making it an excellent text sentiment tool. The advantage of VADER is that it evaluates the mood of any text without prior training of ML models. The result generated by VADER is a dictionary of 4 keys neg, neu, pos and components (compound):
• neg, neu and pos mean negative, neutral and positive respectively. Their sum must be equal to 1 or close to it in a floating point operation.
• Compound corresponds to the sum of the valency scores of each word in the lexicon and determines the degree of mood, and not the actual value, unlike the previous ones. Its value ranges from -1 (the strongest negative mood) to +1 (the strongest positive mood). The use of a composite score may be sufficient to determine the main tone of the text. Compound ≥ 0.05 for positive mood, compound ≤ -0.05 for negative mood, compound ranges from -0.05 to 0.05 for neutral mood
Try Google Colab: https://colab.research.google.com/drive/1_Y7LhR6t0Czsk3UOS3BC7quKDFnULlZG?usp=sharing
Example: https://towardsdatascience.com/social-media-sentiment-analysis-in-python-with-vader-no-training-required-4bc6a21e87b8
Not every classification problem needs machine learning models: sometimes even simple approaches can give excellent results. For example, VADER (Valence Aware Dictionary and sEntiment Reasoner) is a vocabulary and rule based sentiment analysis model. The project source code is available on Github under the MIT license: https://github.com/cjhutto/vaderSentiment
VADER can efficiently handle dictionaries, abbreviations, capital letters, repetitive punctuation marks, emoticons (😢 , 😃 , 😭, etc.), etc., which are commonly used in social networks to express sentiment, making it an excellent text sentiment tool. The advantage of VADER is that it evaluates the mood of any text without prior training of ML models. The result generated by VADER is a dictionary of 4 keys neg, neu, pos and components (compound):
• neg, neu and pos mean negative, neutral and positive respectively. Their sum must be equal to 1 or close to it in a floating point operation.
• Compound corresponds to the sum of the valency scores of each word in the lexicon and determines the degree of mood, and not the actual value, unlike the previous ones. Its value ranges from -1 (the strongest negative mood) to +1 (the strongest positive mood). The use of a composite score may be sufficient to determine the main tone of the text. Compound ≥ 0.05 for positive mood, compound ≤ -0.05 for negative mood, compound ranges from -0.05 to 0.05 for neutral mood
Try Google Colab: https://colab.research.google.com/drive/1_Y7LhR6t0Czsk3UOS3BC7quKDFnULlZG?usp=sharing
Example: https://towardsdatascience.com/social-media-sentiment-analysis-in-python-with-vader-no-training-required-4bc6a21e87b8
GitHub
GitHub - cjhutto/vaderSentiment: VADER Sentiment Analysis. VADER (Valence Aware Dictionary and sEntiment Reasoner) is a lexicon…
VADER Sentiment Analysis. VADER (Valence Aware Dictionary and sEntiment Reasoner) is a lexicon and rule-based sentiment analysis tool that is specifically attuned to sentiments expressed in social ...
Forwarded from Технологии | Нейросети | Боты
Media is too big
VIEW IN TELEGRAM
Новый Нейрофотошоп
Blended Diffusion: Text-driven Editing of Natural Images
"Учитывая входное изображение и маску, Blended Diffusion изменяет замаскированную область в соответствии с текстовым приглашением, не затрагивая незамаскированные области."
Project page : https://omriavrahami.com/blended-diffusion-page
Github :
https://github.com/omriav/blended-diffusion
Blended Diffusion: Text-driven Editing of Natural Images
"Учитывая входное изображение и маску, Blended Diffusion изменяет замаскированную область в соответствии с текстовым приглашением, не затрагивая незамаскированные области."
Project page : https://omriavrahami.com/blended-diffusion-page
Github :
https://github.com/omriav/blended-diffusion
Forwarded from Технологии | Нейросети | Боты
Как создать демонстрацию машинного обучения в 2022 году.
Интересная статья:
https://nicjac.dev/posts/how-to-build-machine-learning-demo-in-2022
Интересная статья:
https://nicjac.dev/posts/how-to-build-machine-learning-demo-in-2022
Forwarded from DL in NLP (Vlad Lialin)
Advanced Topics in MultiModal Machine Learning
cmu-multicomp-lab.github.io/adv-mmml-course/spring2022
Весьма up-to-date курс по мультимодальному обучению от Carnegie Mellon University. В основном обсуждают модальность картинка+текст но говорят немного и про видео. В курсе ней есть как и уже стандартные подходы вроде VL-BERT, так и очень интересный топик по длинным трансформерам и памяти.
Видео нету, но есть очень pdf с очень подробными lecture notes (например вот лекция по длинным трансформерам). Если вы погружаетесь в мультимодальную тему, рекомендую использовать этот курс в качестве гайда. Сам постараюсь почитать.
cmu-multicomp-lab.github.io/adv-mmml-course/spring2022
Весьма up-to-date курс по мультимодальному обучению от Carnegie Mellon University. В основном обсуждают модальность картинка+текст но говорят немного и про видео. В курсе ней есть как и уже стандартные подходы вроде VL-BERT, так и очень интересный топик по длинным трансформерам и памяти.
Видео нету, но есть очень pdf с очень подробными lecture notes (например вот лекция по длинным трансформерам). Если вы погружаетесь в мультимодальную тему, рекомендую использовать этот курс в качестве гайда. Сам постараюсь почитать.
cmu-multicomp-lab.github.io
11-877 AMML
11-877 Advanced Topics in Multimodal Machine Learning - Carnegie Mellon University - Spring 2022
Forwarded from GitHub Community
ImageAI – Библиотека python, созданная для того, чтобы дать разработчикам возможность создавать приложения и различные системы с возможностями автономного компьютерного зрения используя несколько строк кода.
ImageAI поддерживает обнаружение объектов на видео и отслеживание объектов с помощью RetinaNet, YOLOv3 и TinyYOLOv3, обученных набору данных COCO.
GitHub | #Python #AI
ImageAI поддерживает обнаружение объектов на видео и отслеживание объектов с помощью RetinaNet, YOLOv3 и TinyYOLOv3, обученных набору данных COCO.
GitHub | #Python #AI
Forwarded from GitHub Community
pyAudioAnalysis – Библиотека Python для извлечения звуковых функций, классификации, сегментации звуковых данных
С помощью pyAudioAnalysis вы можете классифицировать неизвестные звуки, распознавать звуки с помощью машинного обучения и многое другое..
GitHub | #Python #Audio #Analyzer #Interesting
С помощью pyAudioAnalysis вы можете классифицировать неизвестные звуки, распознавать звуки с помощью машинного обучения и многое другое..
GitHub | #Python #Audio #Analyzer #Interesting
Forwarded from Senior Python Developer
Преобразование текста в речь
Рассмотрим модуль pyttsx3, позволяющий озвучивать текст прямо во время выполнения программы. Для запуска кода с картинки необходимо установить модуль при помощи
Модуль позволяет менять настройки произношения. Полная документация доступна по ссылке: https://pypi.org/project/pyttsx3/
Рассмотрим модуль pyttsx3, позволяющий озвучивать текст прямо во время выполнения программы. Для запуска кода с картинки необходимо установить модуль при помощи
pip install pyttsx3. Запущенная программа спросит, как у вас дела, и скажет, что любит макароны.Модуль позволяет менять настройки произношения. Полная документация доступна по ссылке: https://pypi.org/project/pyttsx3/