I have sat through a lot of AI demos where the face looks perfect but the conversation is unstable.
You say something, it waits. You pause to think, it talks over you. You interrupt, and it finishes its sentence anyway.
Everyone nods and nobody says the obvious thing, which is that this is not a conversation.
So we built Human AI Agents.
What it is
One still photo, a persona written in plain language, and a voice. You get an agent you can talk to and interrupt.
What runs underneath
Two face models behind a single API. Portrait for speed and scale, Presence for expressiveness. Both stream over WebSocket and drop into Pipecat or LiveKit.
What we actually spent the time on
Turn-taking. Knowing when someone has finished a sentence rather than paused to think. It is the unglamorous part and it is most of the product.
Try to break it. Interrupt it mid-sentence, talk over it, trail off, use an accent, most of all have fun!
I'll be around all day.
PH 用户
Congrats on the launch! ✨ Got some nice recommendations from Sofia for travelling and it continues the discussions pretty smoothly when interrupted or when I interject something randomly (and knows to respond to that also). It still is obvious that it's an AI, but the experience is pretty good.
PH 用户
So happy to see Ojin out today after all the hard work! Proud to work with this team 🔥
PH 用户
Proud and honored to be a part of the Ojin team. Hard work do pay off. Congratulations to everyone at Ojin. My favourite is Presence. 🤩
PH 用户
I'm very excited to finally have Ojin released!! A lot of engineering efforts went into making this product possible and I'm so proud of the team. Can't wait for people to try it out!
I have sat through a lot of AI demos where the face looks perfect but the conversation is unstable.
You say something, it waits. You pause to think, it talks over you. You interrupt, and it finishes its sentence anyway.
Everyone nods and nobody says the obvious thing, which is that this is not a conversation.
So we built Human AI Agents.
What it is
One still photo, a persona written in plain language, and a voice. You get an agent you can talk to and interrupt.
What runs underneath
Two face models behind a single API. Portrait for speed and scale, Presence for expressiveness. Both stream over WebSocket and drop into Pipecat or LiveKit.
What we actually spent the time on
Turn-taking. Knowing when someone has finished a sentence rather than paused to think. It is the unglamorous part and it is most of the product.
Try to break it. Interrupt it mid-sentence, talk over it, trail off, use an accent, most of all have fun!
I'll be around all day.