<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Voice-AI :: Category :: Documentation for AI Services</title><link>https://docs.ai.gwdg.de/en/categories/voice-ai/index.html</link><description/><generator>Hugo</generator><language>en</language><atom:link href="https://docs.ai.gwdg.de/en/categories/voice-ai/index.xml" rel="self" type="application/rss+xml"/><item><title>Voice-AI FAQ</title><link>https://docs.ai.gwdg.de/en/user/ai-services/voice-ai/faq/index.html</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://docs.ai.gwdg.de/en/user/ai-services/voice-ai/faq/index.html</guid><description>How to use Voice-live tool effectively? For Best Accuracy, Select Your input Language! While the Auto Detect feature is powerful, the transcription model achieves the highest accuracy and speed when you specify the spoken language beforehand. If you know what language will be spoken, selecting it from the dropdown is highly recommended.
How to do transcription/translation in an online meeting or generally from a system audio? For security and privacy reasons, web browsers cannot directly “listen” to your computer’s speaker output. To transcribe a meeting (from Zoom, Teams, etc.) or any other audio playing on your computer, you need to use a Virtual Audio Cable. This free software creates a virtual “loopback” device that routes your speaker audio to a virtual microphone, which you can then select in this web app.</description></item><item><title>AI Services</title><link>https://docs.ai.gwdg.de/en/user/ai-services/index.html</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://docs.ai.gwdg.de/en/user/ai-services/index.html</guid><description>The AI services offered by GWDG provide a versatile platform for practical AI applications. Our portfolio ranges from a chatbot with advanced capabilities (e.g. retrieval augmented generation, tool integration, MCP) to speech transcription and image processing. All services can be conveniently used via a web interface or seamlessly integrated into existing systems via an OpenAI-compatible API. Developed as part of the KISSKI project (AI Service Centre for Sensitive and Critical Infrastructures), the services meet high data protection requirements and are therefore particularly suitable for sensitive application scenarios.</description></item><item><title>Voice AI</title><link>https://docs.ai.gwdg.de/en/user/ai-services/voice-ai/index.html</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://docs.ai.gwdg.de/en/user/ai-services/voice-ai/index.html</guid><description>One of KISSKI’s standout offerings is its AI-based transcription and captioning service, Voice-AI. Utilizing High-Performance Computing (HPC) infrastructure, Voice-AI leverages the Whisper (large-v2) to transcribe audio and generate video captions swiftly. Trained on 680,000 hours of labeled data, Whisper rivals professional human transcribers in performance, offering reliable automatic speech recognition (ASR) and speech translation across various datasets and domains. Users can choose between tasks such as transcription and translation to suit their needs, and notably, this KISSKI service will be available for free.</description></item></channel></rss>