AI Service
Port: 8084 (HTTP) · 9094 (gRPC)
Database: omninet_ai
Providers: Google Gemini (primary) + Ollama (fallback)
The AI service is the intelligence layer of OmniNet, powering a full-featured, Gemini-like conversational AI experience.
Features
- Multi-Session Chat — Users can have multiple concurrent chat sessions, each with independent history.
- Gemini Integration — Primary AI provider using the Gemini REST API (
gemini-2.0-flash). - Ollama Fallback — Automatic failover to a local Ollama model when Gemini is unavailable.
- SSE Streaming — AI responses stream to the frontend token-by-token via Server-Sent Events.
- Background Persistence — Responses are persisted to the database even if the client disconnects mid-stream.
- Web Search Tool Calling — Gemini can call a web search tool and return source citations.
- Conversation Memory — Full message history is maintained per session and sent as context.
- Voice Input (STT) — Speech-to-text via browser
MediaRecorderAPI. - Markdown Rendering — Responses are rendered with full markdown and syntax-highlighted code blocks.
- Message Editing — Users can edit and resend messages, truncating history to that point.
- Automatic Retry — Exponential backoff on Gemini API failures before failover.
REST API
Chat Sessions
| Method | Path | Description |
|---|---|---|
GET | /api/chat/sessions | List all sessions for user |
POST | /api/chat/sessions | Create a new session |
DELETE | /api/chat/sessions/{id} | Delete a session |
GET | /api/chat/sessions/{id}/messages | Get session message history |
AI Chat
| Method | Path | Description |
|---|---|---|
POST | /api/ai/chat/stream | Send message, receive SSE stream |
POST | /api/ai/chat | Send message, receive full response (non-streaming) |
Streaming Response Format
The /api/ai/chat/stream endpoint uses ResponseBodyEmitter (SSE) and sends events:
data: Hello
data: world
data: !
data: [DONE]The frontend accumulates tokens and renders them progressively.
AI Provider Architecture
interface AiProvider {
String chat(List<Message> history, String userMessage);
void chatStream(List<Message> history, String userMessage, ResponseBodyEmitter emitter);
}Providers:
GeminiAiProvider— Primary. Uses Gemini REST API with function calling.OllamaAiProvider— Fallback. Uses local Ollama REST API.
Failover order:
gemini-2.0-flash(primary model)gemini-1.5-flash-latest(secondary model)- Ollama (
llama3or configured model)
Gemini Tool Calling — Web Search
When the user asks a question that requires current information, Gemini calls the webSearch tool:
{
"name": "webSearch",
"parameters": {
"query": "latest Spring Boot release"
}
}The service executes the search, feeds results back to Gemini, and Gemini generates a response with inline source citations.
Message Turn Alternation
Gemini requires strict alternation of user / model turns in the conversation history. The service enforces this by:
- Filtering out consecutive same-role messages.
- Always starting history with a
userturn. - Appending the current user message as the final
userturn before sending.
Voice Input Pipeline
- Browser records audio via
MediaRecorderAPI. - Audio blob is sent to the AI service.
- AI service transcribes via
SpeechServiceGrpc(Python STT service). - Transcribed text is used as the chat message.
Session Message Cache
The frontend maintains an in-memory session message cache (Map<sessionId, Message[]>). When switching sessions, cached messages are displayed instantly without waiting for the API response. Background fetches update the cache.