| Ticket #32 requirement | Implementation |
|---|---|
| Agent configuration/instructions | /agent-ai and /api/ai/agents; company-scoped agent with explicitly linked ready documents |
| Relevant knowledge retrieval | Compatible model/dimension embeddings, cosine ranking, both agent and document company filters |
| Context | Latest 12 text messages, agent instructions, current question and authorized source chunks |
| Traceable sources | Validated chunk IDs restricted to sources actually sent; full source snapshots stored with drafts |
| Sensitive actions/human approval | Every answer requires operator approval; model output never executes tools or sends directly |
| Handoff and audit | Unsupported factual answers, provider failure or invalid response record a handoff and removes AI control; decisions recorded with actor/time |
The similarity threshold is a retrieval heuristic, not a calibrated confidence probability. Human review is always required, even if the agent configuration would permit automatic operation. Citation validation proves that a referenced source was supplied, not that every generated claim is entailed by it.
The additive ai_reply_drafts and ai_reply_audit tables keep drafts out of the customer message history. After updating the feature branch and using the normal project schema setup, run:
pipenv run flask ai-schema-upgradeThis idempotent command creates only these two tables with SQLAlchemy on SQLite or PostgreSQL; it does not recreate existing tables or delete data. The team integrating the consolidated migration tickets #17/#43 must include these model tables in that migration baseline. No migration history is rewritten here. Back up an existing database using the team's normal procedure before schema changes.
Set in the private backend .env:
AI_SERVICE_URL=https://mac-mini-de-stonetech.taildc6e48.ts.net/v1/agents/customer-service/respond
AI_SERVICE_MODEL=qwen2.5:3b
AI_SERVICE_COMPANY_ID=YOUR_LOCAL_COMPANY_ID
AI_SERVICE_API_KEY=YOUR_PROVISIONED_AGENT_KEY
AI_MIN_SIMILARITY=0.5For multiple companies, use AI_SERVICE_COMPANY_KEYS as a JSON object mapping local company IDs to independently provisioned service credentials. Do not copy another developer's company ID or secrets. No matching key means a safe handoff. Existing KNOWLEDGE_EMBEDDINGS_* settings remain necessary. Generate embeddings with the same model as the stored chunks (embeddinggemma, 768 dimensions in the Mini installation).
The provider request is {message, conversation_id: null, response_format: "clientflow_v1"}. The client rejects a non-null returned conversation ID or an unexpected model. The endpoint has a 12000-character structured input limit (individual customer questions remain limited to 4000 characters): least-relevant complete sources and then older history are dropped; a remaining oversized context is handed to a human. The question and agent instructions are not truncated. Confirm the provider remains stateless for null conversation IDs before deploying real customer data. The client cannot audit the Mini's internal retention implementation.
The browser calls only ClientFlow /api; there are no browser calls to the Mini and no VITE-prefixed AI keys. Configure VITE_BACKEND_URL with that Codespace backend origin (without /api or /login) and restart Vite after editing it. The current application bootstrap still requires this setting; an empty value is not sufficient for the whole application. The API origin must be reachable/authenticated through the Codespace forwarding setup, and FRONTEND_ORIGIN must match the frontend origin. A laptop localhost URL cannot reach a cloud Codespace backend. Each Codespace backend must itself have authorized Tailscale connectivity to the Mini. Having Tailscale on the developer's laptop alone does not put a cloud Codespace in the tailnet. Use your organization's approved Tailscale setup inside the Codespace (or an authorized private gateway) and verify there:
curl --max-time 15 -sS https://mac-mini-de-stonetech.taildc6e48.ts.net/v1/health/readyExpect status: ready. Provision the appropriate company credential or a ClientFlow platform credential for stateless mode, prepare the two tables, then restart Flask. If connectivity is absent, the rest of ClientFlow works and AI drafts hand off to a human. Never make the Mini public or disable authentication to work around connectivity.
- In Conocimiento, upload and process a document.
- In Agentes IA, create an agent and select its processed company documents (owner/admin/manager).
- In Conversaciones, select an agent and click Asignar.
- Enter the customer's question in Asistente IA and generate a draft.
- Inspect cited source excerpts. Approve or reject the draft.
- Approval creates one outbound
storedmessage that the authenticated web chat receives through its normal polling endpoint. It does not send external WhatsApp/email messages. - New messages, revoked/changed sources, changed agent instructions or human takeover invalidate old draft approval.
All routes require a bearer token and X-Company-ID with active membership/subscription. Technician accounts cannot use AI operations. Owner/admin/manager configure agents; owner/admin/manager/agent generate and review.
- GET/POST
/api/ai/agents: list or create/update configuration and linked documents. - GET/POST
/api/conversations/<id>/ai-drafts: last 20 drafts or generate with{question}. - POST
/api/ai/drafts/<id>/decision:{action: "approve"}orreject. - GET
/api/ai/drafts/<id>/audit: ordered actor/action/time records.
Run backend tests using a disposable database, never the application database:
AUTH_TEST_DATABASE_URL=sqlite:// PYTHONPATH=src:tests pipenv run python -m unittest discover -s tests -p 'test_*.py'
node --test tests/frontend/calendar.test.mjs
npm run buildCoverage includes tenant/document assignment isolation, role restrictions, follow-up context, provider contract/errors, source validation, clarification without sources, low-relevance evidence exclusion, revoked sources, stale conversations, duplicate decisions and audit persistence. SQLite is used for automated database tests; PostgreSQL concurrency and teammates' private Codespaces require validation in those environments.
For a ClientFlow-owned inference service that does not load or retain conversational memory, set AI_SERVICE_AUTH_MODE=platform_stateless and provision AI_SERVICE_PLATFORM_KEY on the backend. New companies need no ID mapping. The default company mode retains separate credentials and fails closed for unmapped IDs.
The platform key authenticates the application, not an end user. All AI routes still require an active company membership, eligible subscription and operator role. Document joins, conversation history, agent selection and draft approval remain company scoped. The request contains only authorized context and conversation_id=null; a response returning remote conversation memory is rejected. Never expose the key to the browser. Do not use this mode with a service that retrieves shared tenant data.
The Mini agent source shown in the setup task sends only a fixed system prompt and the current request to Ollama, without fetching conversation memory. Infrastructure logging/retention is a separate deployment concern and has not been audited here.
Standalone greetings use localized resource text after company/agent authorization. Other requests are generated with recent conversation history. Retrieval searches both current input and recent customer context so an earlier greeting does not suppress relevant evidence. Authorized adjacent excerpts can accompany a relevant excerpt from the same document. Unrelated documents are not added.
The Mini must support clientflow_v1: structured JSON generation, native conversation
turns, and a 12000-character request budget. It derives needs_human from response_type
and validates the output. Do not send large string repetition bounds to Ollama's grammar
compiler; enforce the 6000-character output limit in the API validator instead.
Legacy plain text requests retain their 4000-character input limit.
Clarifications may request missing customer details without citations. Factual answers still require authorized source IDs. Source checks do not prove the truth of a generated sentence; operators must review wording and claims, including apparent scheduling promises. No draft is sent automatically. A question mark is not mandatory for polite requests.
The UI displays the specific failure reason; historical failed drafts remain visible. A successful new draft does not delete the audit history. The latest-message shortcut uses only messages on the currently displayed history page.
The browser acceptance run used a disposable SQLite database and a synthetic account, with real Mini inference. A visitor requested a wardrobe replacement, an operator generated and approved the reply, and the visitor received it. A follow-up mentioning sliding doors and Triana produced a question about measurements; its approved reply also reached the visitor. No real customer messages were sent in this run.
The Mini changes are already running but belong to a separate repository. They are not committed by this ClientFlow PR. The temporary SSH authorization was removed; recording that remote commit requires restoring authorized repository access.