Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.
-
Updated
Jul 17, 2026 - Python
Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.
A curated list of full-duplex spoken dialogue models & benchmarks
Jarvis × Codex v0.2.0:macOS 透明全息语音助手,本机“嗨 Jarvis”唤醒,直连 Codex Voice,在同一线程中对话、执行任务并支持 STOP 中断。
PersonaPlex on Apple Silicon: an MLX port of NVIDIA’s full-duplex speech-to-speech model with realtime local/web modes and offline WAV inference.
[EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decision machine, with a curated list of models, datasets, and benchmarks.
Peer-to-peer, secure, TCP/UDP port forwarding using HTTP(s) relay for NAT/firewall traversal
Self-hosted realtime voice gateway for Hermes Agent. Keep talking while Hermes runs background tasks through local, Gemini Live, and OpenAI Realtime adapters.
Adaptive filters for GNU Radio
WebView attachable full-duplex asynchronous interoperable independent messaging library between .NET and JavaScript.
Linux kernel multi- bidirectional- logical- channel communication driver for connecting User/Kernel space entities via full-duplex symmetrical transport.
Game-Time: Evaluating Temporal Dynamics in Full-Duplex Spoken Language Models
Full-duplex voice plugin for DeepSeek Harness: local zipformer2 streaming ASR (no API key) → editable draft; Edge TTS or local VITS / Kokoro read-aloud with live captions; true barge-in; hardened HTTP surface + model SHA256 pinning; compatible with dsh 0.1.1 → 0.1.2. · DSH 语音双工插件:流式识别入草稿、按句朗读+实时字幕、开口即打断;免 API Key、安全加固、新旧版本兼容。
Real-time, full-duplex AI voice bot integrating NVIDIA's PersonaPlex with Twilio Media Streams for natural speech-to-speech conversations.
Simplest and fully-customizable RPC standalone infrastructure.
Real-time serving of full-duplex interaction models: per-tick deadline scheduling + KV-budget admission control. Systems research.
600ms-budget full-duplex voice agent — cascade ASR·RAG·LLM·TTS (458ms) + end-to-end Omni (250ms) | Voice · ASR · TTS · LLM · Infra · Posttraining
Real-time streaming talking-head avatar: PCM audio in, MuseTalk lip-synced video out over a self-developed WebSocket transport. Full-duplex voice sessions with pluggable ASR / LLM / TTS spokes and ms-level barge-in.
A robust, asynchronous WebSocket client built from scratch using C++17, Winsock, and OpenSSL. Implements the RFC 6455 handshake, bitwise frame masking, and secure WSS communication without external WebSocket libraries.
🚀 With the aid of a go named pipe to achieve fast communication with the process(Full Duplex)
To associate your repository with the full-duplex topic, visit your repo's landing page and select "manage topics."