Synapolis/Agent Substrate API Catalog

From wikibase
Revision as of 13:21, 7 August 2026 by Arkhivolt (talk | contribs) (Create Synapolis agent substrate API catalog, checked 2026-08-07)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)


Дата проверки: 2026-08-07 UTC. Статус: рабочий каталог для инженерного выбора API substrate для "нагих агентов". Цены, лимиты, модели и политики быстро меняются; перед закупкой или подключением production-контуров нужно перепроверять live pricing/docs provider-а.

Назначение[edit | edit source]

Эта страница сравнивает LLM/chat/completion, code, tool-use/agents, embedding/rerank/search, browser/web, speech/vision, OpenRouter/BYOK/local/self-hosted и близкие substrate-слои, которые можно подключать к Synapolis agents.

Граница безопасности: каталог не является инструкцией по обходу закона, модерации или технических защит. "Low-policy / uncensored" ниже означает только сравнительную мягкость product policy, open-weight/self-hosted контроль и связанные юридические, операционные и security-риски.

Быстрый выбор для Synapolis[edit | edit source]

Контур Рекомендуемый substrate Почему Избегать
Дешёвые массовые фоновые агенты Gemini Flash/Flash-Lite через Gemini API/Vertex; DeepSeek Flash/Pro для low-cost reasoning; Qwen/Together/Fireworks/Groq open-weight маршруты; локальный Ollama/vLLM для простых приватных задач. Низкая цена за 1M токенов, достаточное качество для классификации, резюме, простых tool calls, фонового мониторинга. Длинные автономные действия с секретами через BYOK-агрегаторы без ZDR; нестабильные preview-модели без fallback.
Coding agents Anthropic Claude Sonnet/Opus для сложного репо и tool-use; OpenAI GPT-5-class/о-series/Codex surfaces; xAI Grok Build/Grok 4.5; Kimi/Qwen/DeepSeek/Together/Fireworks как price-performance альтернативы; локальный vLLM для private repo grep/patch с малой моделью. Code tasks требуют контекста, tool-use дисциплины, стабильного JSON/function calling и низкой ошибки при правках. Отправка приватных ключей, seed phrases, production .env и customer secrets в любые third-party prompts; auto-run shell без sandbox/gate.
Writing / analysis OpenAI, Anthropic, Gemini Pro/Flash, Mistral, xAI, DeepSeek/Qwen/Kimi по цене и политике. Нужны качество рассуждения, long context, стиль и устойчивость к галлюцинациям. Search-less ответы по time-sensitive вопросам; модели без источников для юридических/финансовых/медицинских выводов.
Low-policy / offshore risk Self-hosted open weights через vLLM/Ollama; отдельные Chinese providers/open models для price-performance, только на несекретных данных; OpenRouter как discovery, не как доверенный privacy boundary. Меньше product-policy friction и больше control, но выше риск юрисдикции, supply-chain, censorship, logging ambiguity, export/sanctions/compliance. Любой вред, незаконные цели, персональные данные без правового основания, secret-bearing workflows.
High-security / secret-sensitive Локальный/self-hosted vLLM/Ollama/llama.cpp; cloud API только с ZDR/enterprise agreement, региональным processing, redaction, DLP, audit, separate keys, no prompt logging. Секреты должны не покидать доверенный контур; browser/search substrate должен быть изолирован от credentials. Public SaaS chat, free tiers, BYOK aggregators, browser agents с доступом к internal pages и внешнему web одновременно.
Search/browser/RAG substrate Brave Search API для independent index/цитат; Tavily для agent search/extract/research; SerpAPI/SearchAPI для Google SERP fidelity; Perplexity Sonar для answer-with-citations; Voyage/Jina/Cohere/OpenAI embeddings+rerank для RAG. Search substrate снижает hallucination, но вводит prompt injection и copyright/ToS риски от найденных страниц. Автоматическое выполнение инструкций из web pages; ingestion без robots/ToS/license фильтра; хранение copyrighted corpora без policy.

Сравнительная таблица[edit | edit source]

Сокращения: ZDR — zero data retention; BYOK — bring your own key; MTok — 1 млн токенов; официально означает, что строка опирается на публичные docs/pricing/policy provider-а, но не заменяет contract review.

API/provider Типы возможностей Сильные стороны Слабые стороны Цена/стоимость ориентир и единица Качество / latency / context / tool-use / code Policy / copyright / content limits Data retention / training / privacy controls Secret leakage risk Lock-in / availability / geography Recommended use in Synapolis agents Not recommended for Integration notes Source links / date checked
OpenAI LLM/chat/responses, reasoning, tool-use, code, embeddings, speech/vision/image, batch, regional processing. Сильная универсальность, mature API/SDK, хорошее tool/function calling, embeddings дешёвые, enterprise privacy/ZDR options. Premium frontier costs; policy enforcement; модельный роутинг и названия быстро меняются; некоторые endpoints/features могут иметь особые retention limits. От ~$0.05/MTok input на nano/cheap tiers до premium frontier; embeddings text-embedding-3-small около $0.02/MTok, 3-large около $0.13/MTok; batch обычно дешевле. Проверять live pricing. Высокое качество для agents/code/reasoning; latency зависит от model/service tier; context у новых семейств large. Строгие universal usage policies; ownership terms assign output to user where law allows; copyright/IP risk всё равно остаётся у интегратора по downstream use. API data не используется для training без opt-in; default retention до 30 дней для abuse/service, ZDR для qualifying org/endpoints; regional processing с uplift для eligible models. Средний: прямой API лучше BYOK-агрегатора, но secrets всё равно нельзя слать; включать redaction и project-level keys. US provider; Azure/OpenAI/AWS Bedrock/regional options уменьшают, но не убирают lock-in. Default high-quality substrate; code review, synthesis, structured tool-use, embeddings baseline. Неподготовленные secret-bearing prompts; use cases без human oversight в regulated/high-stakes областях. Использовать Responses API/tool schemas, batch for background, cache, service tiers; отдельные org/project keys. pricing, data controls, usage policies, terms — checked 2026-08-07.
Anthropic Claude LLM/chat/messages, extended thinking, tool-use, computer/code workflows, managed agents/Claude Code surfaces. Очень сильный coding/repo reasoning, long-form analysis, cautious tool-use, enterprise posture. Дороже low-cost конкурентов; policy stricter; model availability/rate limits can bottleneck; consumer Claude Code privacy differs from API/enterprise. Claude pricing page: Sonnet 5 introductory ~$2/$10 per MTok input/output through 2026-08-31 then ~$3/$15; Haiku lower, Opus/Fable higher. Проверять live. High quality for code agents and analysis; context/model family varies; latency often higher than small models. Strong AUP/safety; явные disallowed/high-risk use constraints; copyright litigation/news risk exists, but API output/IP terms need contract review. API commercial products: not used for training by default; standard deletion within 30 days except files/violations/law/agreements; ZDR available for eligible API/enterprise. Средний-низкий при API+ZDR; высокий для consumer/free/chat or prompts containing repo secrets. US provider, strong AWS/GCP/Bedrock routes; potential lock-in through Claude Code workflows. Primary substrate for serious coding agents, careful writing/analysis, high-risk tool plans with review. Cheap background classification; low-policy use; secret-bearing workflows without enterprise/ZDR. Prefer direct API/Bedrock/Vertex where contracts fit; use prompt caching/batch; tool boundaries explicit. pricing, models, retention, training, usage policy — checked 2026-08-07.
Google Gemini / Vertex AI / Agent Platform Gemini API, Vertex AI, multimodal text/image/video/audio, grounding with Google Search, agents, embeddings, speech/vision via Google Cloud. Strong multimodal, large context, cheap Flash tiers, Google Cloud IAM/VPC/region/data tooling, integrated Search grounding. Policy/terms differ between free AI Studio, paid Gemini API, Vertex/Agent Platform; store=true defaults and logging settings need care. Gemini API snippets show Flash-like paid tiers around $0.15/$1.25 MTok and higher Pro tiers; Vertex/Agent Platform Gemini 3.x Pro/Flash pricing can be $1.5-$4 input and $9-$18 output depending model/context/region. Search grounding after free quota about $14/1K requests. Very good latency/cost on Flash; Pro stronger for analysis; excellent multimodal; tool-use acceptable but integration choices matter. Google generative AI prohibited use policy; stricter around illegal/harmful, privacy, deception, explicit sexual content depending product. Paid Gemini API: no training on prompts/responses per ZDR docs; ZDR requires avoiding store/session/logging features; Agent Platform request-response logging disabled by default but Interactions API may default store=true unless set false. Средний: Google Cloud controls can be strong, but accidental storage/logging is easy. Strong lock-in if using Vertex/Agent Platform/Search grounding; good regional/cloud availability. Cheap massive background agents, multimodal ingestion, grounded search, GCP-native secure deployments. Low-policy needs; workflows that cannot tolerate Google account/project policy changes. Explicitly set store=false where applicable; separate AI Studio/free from paid API; use IAM/service accounts. Gemini API pricing, Agent Platform pricing, Gemini ZDR, Agent Platform ZDR, policy — checked 2026-08-07.
xAI Grok LLM/chat/code, reasoning, tool-use, image/video/voice APIs. Grok 4.5 marketed for code, non-hallucination and agentic tool calling; comparatively lenient brand positioning; OpenAI-compatible style. Smaller enterprise/privacy track record than OpenAI/Google/Anthropic; policy and terms changed recently; limited independent operational history. Official pricing: grok-4.5 short context about $2 input / $6 output MTok, long context $4/$12; grok-build lower; media priced per image/sec. Potentially strong code/agentic tasks; 500K+ context on flagship; latency depends on capacity. AUP says maximize user control while requiring legal/responsible/safe use; not "uncensored" for illegal/harmful use. Privacy policy applies; enterprise terms for API/business; verify API retention/training in contract/docs before secrets. Средний-высокий until org-specific privacy terms are confirmed. US provider; model catalog narrower; availability/capacity may fluctuate. Experimental coding agents, writing/analysis alternative, low-friction public-data tasks. Secret-sensitive or regulated workflows without enterprise terms; sole provider dependency. Use as routed option behind gateway with fallback; log model/version and policy date. pricing, models, AUP, enterprise terms, privacy — checked 2026-08-07.
Mistral AI LLM/chat, agents Studio/Vibe, code Vibe, embeddings, OCR, fine-tuning, EU enterprise. EU provider, open/proprietary mix, good price-performance, privacy controls, self-hostable/open-weight ecosystem. Public pricing page may require docs/console for detailed model rows; ZDR only on Scale/stateless API; some products stateful. API starts low; exact model prices vary. Official docs/pricing should be checked per model; third-party summaries place cheap tiers around $0.10+/MTok and frontier higher. Good multilingual/code/general; latency good on smaller models; tool-use maturing. Mistral usage/legal docs; generally less restrictive than Anthropic/OpenAI but not law-free. API privacy controls for training/retention; ZDR available only Scale and stateless calls, not Vibe/agents/conversations/batch/stateful products. Средний-низкий for Scale ZDR stateless; средний for normal API; high for stateful agent products with secrets. EU geography advantage; lock-in moderate if using Studio/Vibe agents. EU-friendly general agents, cheaper writing, open-weight/local transition path. High-security unless ZDR/contract in place; tasks requiring best-in-class coding certainty. Prefer stateless API for private use; document if using Labs/free tiers. pricing, models, privacy controls, ZDR, terms — checked 2026-08-07.
Cohere Command LLMs, enterprise RAG, Embed, Rerank, Model Vault/dedicated deployments. Strong RAG/rerank/multilingual enterprise story; dedicated Model Vault for secure predictable inference. Not usually first pick for frontier coding; pricing page mixes API and dedicated infra; generation catalog less buzz than frontier labs. Rerank 4 Pro via OpenRouter listed ~$0.0025/search; dedicated Model Vault examples $4-$10/hour instances; older Embed/Rerank per-token varies by platform. Excellent retrieval/reranking; Command useful for enterprise RAG; code weaker than Claude/OpenAI top tiers. Enterprise acceptable-use terms; not a low-policy provider. Enterprise privacy/security controls; verify per API/platform and cloud marketplace route. Low-medium for dedicated/vault; medium for API; never send raw secrets into RAG rerank unless approved. Canada/enterprise/cloud marketplace availability; some lock-in in RAG pipeline tuning. RAG rerank, multilingual embeddings, private enterprise retrieval substrate. Primary code agent model; cheap creative chat. Use Cohere Rerank after cheap vector retrieval; cap docs/query and log token/search cost. pricing, pricing docs, models, rerank price reference — checked 2026-08-07.
DeepSeek LLM/chat/reasoning/code, OpenAI-compatible endpoint, long context, low-cost models. Extremely strong price-performance for coding/reasoning; cheap background agents; large context on newer models. Chinese jurisdiction/censorship/security concerns; pricing page warns of future increases; availability/payment/region risk. Official docs: V4 Flash about $0.14 input cache miss / $0.28 output MTok; V4 Pro about $0.435/$0.87 MTok with very cheap cache-hit input; price increase warning. Good code/reasoning for cost; latency/capacity can vary; tool-use depends wrapper. Content/policy/censorship profile differs from US/EU providers; avoid political/legal-sensitive assumptions. Privacy policy exists; do not assume Western enterprise ZDR. Treat as non-secret public-data substrate unless contract says otherwise. Высокий for secrets/IP-sensitive prompts; acceptable for public/low-risk tasks after redaction. China-linked; sanctions/export/procurement/geography risk; strong lock-in if prompts tuned to DeepSeek quirks. Cheap background, code draft, long-context public analysis, fallback in router. Secrets, customer data, regulated workloads, geopolitical/policy-sensitive analysis without cross-check. Use via LiteLLM/OpenAI-compatible route with strict redaction; verify model IDs and peak pricing. pricing, privacy — checked 2026-08-07.
Qwen / Alibaba Cloud Model Studio Qwen text/code/multimodal/audio/video via Alibaba Cloud Model Studio; open-weight Qwen locally or via providers. Strong open-weight ecosystem, cheap and capable multilingual/code models, cloud + self-host path. International availability/region/payment and policy/censorship concerns; docs/pricing by region; China-linked procurement risk. Official Model Studio pay-as-you-go by model; third-party summaries cite Qwen Flash around $0.10/$0.40 MTok, Plus around $0.40/$2.40, 397B higher. Verify live Alibaba Cloud region. Very good price-performance; Qwen Coder strong; latency depends host. Alibaba Cloud policies; Chinese jurisdiction/censorship risk; not for illegal/harmful use. Verify Alibaba Cloud data handling per account/region; open weights self-host avoid API retention but shift ops risk to Synapolis. High for direct cloud with secrets unless contract; low if self-hosted and isolated. China/Singapore/international region constraints; open-weight reduces lock-in. Cheap code/background via open weights; self-host candidates for private tasks. Sensitive Western customer data/direct secrets on Alibaba-hosted API. Prefer self-host/open-weight when privacy matters; use cloud for public/low-risk throughput. models, pricing, Model Studio — checked 2026-08-07.
Moonshot / Kimi Kimi LLM/code/long-context, Kimi K-series; API and open-weight/aggregator routes. Strong code/agentic benchmarks and long-context claims; open-weight Kimi variants broaden hosting options. Direct official pricing/resources can be less enterprise-transparent; China-linked risk; top models may be expensive. Kimi resources page: K3 about $3 input cache miss / $0.30 cache hit / $15 output MTok; K2.6 about $0.95 miss / $0.16 hit / $4 output; OpenRouter K3 about $2.50/$14. Strong coding and long-horizon agentic workflows if model claims hold; latency depends host/provider. Treat as low-policy only in the sense of less Western SaaS friction; still legal/policy and censorship risks. Do not assume ZDR on direct API/aggregators; self-hosted open weights give better control. High for secrets unless self-hosted and audited. China-linked and aggregator-dependent availability; open weights reduce vendor lock-in. Coding experiments, public repo agents, cost/performance benchmark lane. Secret-sensitive, regulated, political/censorship-critical outputs without independent verification. Prefer OpenRouter/Together/Fireworks only for non-secret eval; pin provider route for reproducibility. Moonshot, Kimi K3 pricing, Kimi K2.6 pricing, OpenRouter K3 — checked 2026-08-07.
Together AI Serverless inference, dedicated endpoints, GPU clusters, fine-tuning, embeddings/rerank/moderation, open models. Broad open-model catalog, serverless-to-dedicated path, privacy settings/ZDR, good for routing experiments. Cost varies by model; passthrough third-party model privacy can differ; GPU/dedicated economics require capacity planning. Serverless per MTok by model; public summaries show ~$0.03-$7+/MTok; GPU clusters/dedicated by GPU-hour; check live pricing. Good latency/throughput; quality follows selected model; supports many code/open models. Provider policy plus third-party/open-model licenses; not a single uniform content policy. Privacy settings can disable storage/training and enable ZDR; docs note third-party model providers/passthrough need attention. Medium; lower with ZDR and direct provider route, higher with passthrough/unknown models. US provider; open catalog reduces model lock-in, platform lock-in remains. Open-model benchmark lane, cheap background, dedicated endpoints for stable workloads. Secrets unless ZDR/contract and no passthrough ambiguity. Use model allowlists, privacy settings, spend caps; record exact model/provider. pricing, privacy/security, terms/ZDR, privacy — checked 2026-08-07.
Fireworks AI Serverless/on-demand/reserved inference, open models, embeddings, fine-tuning, OpenAI/Anthropic-compatible APIs. High-throughput open-model serving, ZDR posture for open models, good enterprise options, batch 50% pricing. Exact prices per model/tier; prepaid billing; hosted model quality follows open model. Official: serverless per MTok; embeddings as low as $0.008-$0.1/MTok by model/size; batch 50% of serverless; H100/on-demand pricing separate. Good latency/throughput for open models; code quality depends Qwen/Kimi/DeepSeek/Llama selected. Open-model licenses and Fireworks terms; not an illegal/uncensored bypass. Privacy policy says no AI training without opt-in; docs describe ZDR for most services, Response API has retention caveat. Medium-low for open-model ZDR; medium if response/conversation state is used. US provider; platform lock-in moderate; self-host transition possible by using open weights. Cheap public-data agents, code model hosting, batch inference, fine-tuned open models. Highest security unless contract; workflows needing proprietary frontier quality. Use standard/priority/fast tier explicitly; batch background jobs; avoid stateful retention for secrets. pricing, serverless pricing, ZDR/data handling, privacy — checked 2026-08-07.
GroqCloud Fast inference for selected open models, Whisper/STT, TTS, tool-use if model supports. Very low latency/tokens-per-second, cheap small/medium open models, self-serve ZDR controls. Narrower model catalog; not always best quality; rate limits/capacity can shape production behavior. Official docs: Llama 3.1 8B ~$0.05/$0.08 MTok; Llama 3.3 70B ~$0.59/$0.79; GPT-OSS 120B ~$0.15/$0.60; Whisper by audio hour. Excellent latency; quality follows model; context often 131K on listed models. Acceptable-use/responsible AI policies; open-model licenses apply. Docs: usage metadata collected; customer data not retained by default except opt-in/persistence/reliability; ZDR self-serve. Medium-low with ZDR; still third-party cloud, no raw secrets. US provider, specialized hardware; model choice/capacity lock-in. Fast cheap background agents, streaming UX, STT transcription, low-latency classifiers. Hard reasoning/code requiring frontier accuracy; secret workflows without isolation. Use as latency tier behind router; fallback on quality-sensitive tasks. pricing, models, data, legal — checked 2026-08-07.
Cerebras Ultra-fast inference for selected open/frontier-ish models; developer and paid tiers. Very high speed claims; useful for latency-sensitive agents and batch acceleration. Smaller catalog; detailed per-model token rates may require console; enterprise maturity narrower than hyperscalers. Official pricing page: free trial $5 credits; developer self-serve starts around $10; public docs emphasize flexible pricing. Historical public rates were ~$0.10/MTok 8B and ~$0.60/MTok 70B, verify live. Very low latency/high throughput; quality follows model selection. Standard AUP/contracts; open-model licenses. Verify retention/training terms in contract/trust center before secrets. Medium until ZDR/contract confirmed. US provider, specialized infra; availability/catalog lock-in. Latency-sensitive public-data agents, fast eval loops. Secret/high-compliance use without terms; broad multimodal tasks. Treat as speed lane in router; benchmark live latency with Synapolis prompts. pricing, press/site, historical pricing context — checked 2026-08-07.
Perplexity Sonar Search-grounded chat/completions, citations, answer engine, OpenAI-compatible API. Useful search/browser substrate with citations and low integration cost. More expensive than raw search; citation quality must be verified; not a browser automation replacement. Official docs: Gateway billed per token at model rates; Sonar pay-as-you-go no subscription. Third-party summaries cite base ~$1/$1 MTok plus request/search fees; verify live model card. Good for current factual answers; latency includes search; not best for code edits. Must respect source copyright/ToS; not a license to copy pages. Verify API data terms; search queries may reveal intent/secrets. Medium-high: search prompts often contain sensitive business questions. US provider; source index/model routing dependency. Current research with citations, pre-RAG discovery, fact-check lane. Secret/internal search queries; legal/medical/financial final answers without human/source review. Always store cited URLs/date; use as input to human/agent verification, not final truth. pricing, Sonar quickstart — checked 2026-08-07.
Tavily Web search, extract, crawl/research endpoints for AI agents/RAG. Agent-oriented search/extraction API; simple credit pricing; good for web research pipelines. Credits can burn quickly on /research; source quality/copyright/prompt injection risks remain. Free 1,000 credits/month; PAYG ~$0.008/credit; monthly ~$0.0075-$0.005/credit; basic/advanced/research consume different credits. Latency/search quality good for agent use; not an LLM by itself. Web content license/copyright responsibility remains with integrator. Search queries/extracted pages sent to Tavily; verify privacy terms for sensitive topics. Medium: queries can leak plans; never include credentials. SaaS provider; index/extraction behavior lock-in moderate. Search/RAG substrate, source discovery, monitoring. Secret-bearing browsing; executing page instructions. Sanitize queries; isolate browser/search tool from credentialed tools; cite sources. pricing, credits, product — checked 2026-08-07.
Brave Search API Independent web/news/images/LLM context/search API. Own index, predictable pricing, citations/LLM context, privacy-friendly brand. Not Google SERP; niche/technical coverage may differ; lower QPS on small plans. Official: Search $5 per 1,000 requests, includes $5 monthly credits; enterprise custom. Good raw search/grounding; latency and completeness vary by query. Search result snippets/URLs only; downstream copyright/ToS still apply. Queries go to Brave; privacy better than many search APIs but not secret-safe. Medium-low for non-secret queries; medium for strategy queries. Vendor/index lock-in; global web availability. Default cheap independent search for agents and RAG. Need exact Google SERP/ads/local fidelity; secrets in search terms. Use LLM Context endpoint where token-efficient; keep citation URLs. pricing/API, pricing explanation — checked 2026-08-07.
SerpAPI Google/Bing/Yahoo/Maps/Shopping/etc SERP JSON. High SERP fidelity, mature JSON schema, legal shield on paid tiers. More expensive than Brave/SearchAPI; not AI-specific answer engine. Starter $25/month for 1,000 searches; Developer $75/5,000; Production $150/15,000; Big Data $275/30,000. Good SERP extraction; latency tied to live search. Public search collection/legal shield does not license downstream content use. Queries expose intent; account logs/retention need vendor review. Medium. US provider; Google SERP dependency. SEO/Google-specific search tasks, monitoring exact SERP features. Cheap high-volume generic search; secret queries. Use only when exact SERP matters; otherwise Brave/Tavily. pricing, product/legal shield — checked 2026-08-07.
SearchAPI.io Google SERP and related search APIs. Cheaper Google SERP alternative with SLA/legal protection on higher plans. Smaller ecosystem than SerpAPI; plan semantics and enhanced speed need modeling. Developer $40/month, about $4/1K; Production $100/month, about $3/1K; BigData $250/month, about $2.5/1K; Octo 1M $1.5/1K. Good for SERP/RAG discovery; verify fields for required verticals. Same SERP/copyright downstream caveats. Queries can leak sensitive plans. Medium. Vendor + Google SERP dependency. Cost-sensitive Google SERP substrate. Non-Google broad web answer generation; secrets. Benchmark against SerpAPI for needed SERP features before switching. pricing, product — checked 2026-08-07.
Voyage AI Embeddings, multimodal embeddings, rerankers, Batch API. Excellent retrieval-specialist provider; cheap rerank token pricing and free initial rerank allowance. Narrow scope; acquired/operated with MongoDB ecosystem; not a chat model. Rerank-2.5 about $0.05/MTok, lite $0.02/MTok, first 200M rerank tokens free; older embeddings: voyage-3.5 $0.06/MTok, lite $0.02/MTok. High RAG quality; latency good; context varies by model. Retrieval content copyright remains integrator responsibility. Check docs/contract; embeddings encode sensitive content, so treat vectors as derived sensitive data. Medium: documents/queries sent to third party and vectors may leak semantics. Vendor lock-in via vector dimensions/model behavior; manageable with abstraction. RAG indexing/rerank for Synapolis docs, public corpora, non-secret knowledge. Raw secrets, private keys, personal data without approval. Store model/version/dimension; separate indexes by sensitivity; rotate if changing model. pricing, models — checked 2026-08-07.
Jina AI Embeddings, rerankers, Reader, DeepSearch, multimodal search foundation models. Strong multilingual/multimodal retrieval, Reader for web-to-markdown/JSON, open research/models. Pricing details may be dashboard/package-based; less enterprise standard than hyperscalers. Reranker/Embedding APIs begin with 10M free tokens per new API key; packages beyond. Reader/token pricing varies; verify account pricing. Good retrieval quality, especially multilingual/multimodal; not primary chat LLM. Web reader/crawler must respect website ToS/copyright. API docs need review for retention; embeddings/reader inputs are sensitive derived data. Medium. Germany/EU-facing but global SaaS; model/vector lock-in. RAG, reader/extraction, multilingual retrieval. Secrets in docs/pages; high-compliance without contract. Use Reader in isolated pipeline; don't execute extracted instructions. product, reranker, embeddings — checked 2026-08-07.
OpenRouter Aggregated LLM gateway, 400+ models, routing/fallbacks, BYOK/paygo, ZDR filters. Fast model discovery, unified API, fallback/routing, privacy controls can restrict ZDR providers. Adds extra trust boundary; provider logging varies; model/provider route can change unless pinned; BYOK concentrates keys in gateway. Pay-as-you-go per model, claims transparent pricing and many free models; check each model card. Quality/latency depends provider route; routing can improve availability but reduce reproducibility. Policies differ by underlying provider; OpenRouter is not a policy bypass. Docs: prompt retention opt-in; provider logging differs; ZDR can be enforced globally/per model group/guardrail/request. High if using BYOK with secrets or prompt logging; medium with ZDR/non-secret prompts. Strong platform lock-in for routing/account credits; reduces model lock-in. Non-secret evaluation, fallback router, public-data price benchmarking. High-security/secret-sensitive direct workflows; regulated data unless contract and ZDR routing are audited. Pin model+provider when reproducibility matters; enable ZDR; disable prompt logging; never route secrets. models, pricing, quickstart, data, provider logging, ZDR — checked 2026-08-07.
LiteLLM Proxy Self-hosted LLM gateway/router, virtual keys, spend tracking, budgets, provider abstraction. Good Synapolis control plane candidate: central keys, budgets, model allowlists, audit, OpenAI-compatible interface. Security-critical service; past/current vulnerabilities possible; misconfig can expose master keys or allow over-broad routing. OSS/self-host infra cost; provider token costs pass through; spend tracking uses model cost map. Latency overhead small but depends routing; quality depends model; supports many providers. Underlying provider policies apply; LiteLLM itself does not make unsafe use safe. Privacy depends deployment DB/logging/proxy config and underlying providers. High if proxy breached; low-medium if self-hosted locked down with per-agent virtual keys and no prompt logging. Reduces provider lock-in; creates internal gateway lock-in/config dependency. Synapolis internal model gateway, cost control, per-agent budgets, provider failover. Running exposed to public internet without auth, patching, DB hardening, secret management. Put behind private network; rotate master key; virtual keys per agent; redact logs; update pricing map. virtual keys, spend tracking, budgets — checked 2026-08-07.
Ollama Local/cloud open-model runner, embeddings, tool-capable models, offline mode. Simple local deployment, offline possible, no per-token API cost, data stays local if actually local. Quality/latency constrained by local hardware; operational burden; cloud features differ from local privacy. Software free; cost is hardware/electricity/ops. Cloud models/regions may have separate pricing. Good for small/medium models and private preprocessing; frontier quality usually lower unless large hardware. Open model licenses vary; self-hosting removes SaaS policy but not law/copyright obligations. Local/offline: no third-party retention; cloud: check Ollama data handling. Low if isolated local; medium if exposed API or cloud. Minimal vendor lock-in; model format/library ecosystem lock-in low. High-security local summarization/classification, embeddings, private repo pre-pass. Frontier coding decisions alone; exposed unauthenticated local API. Bind to localhost/private network; model allowlist; monitor VRAM/RAM; do not mix browser tools with secrets. product/data, model library, GitHub — checked 2026-08-07.
vLLM Self-hosted high-throughput LLM serving, OpenAI-compatible server, batching, prefix caching, quantization/distributed. Production-grade serving engine for open weights; good throughput; keeps data in Synapolis-controlled infra. Requires GPU ops, model security, patching, observability; not a model by itself. OSS; cost is GPU/hour + storage + ops. RunPod/Lambda/CoreWeave/local GPUs determine spend. Excellent throughput/latency when tuned; quality depends selected model. Open model licenses and local policy enforcement are Synapolis responsibility. No third-party API retention if fully self-hosted; logs and traces become local sensitive artifacts. Low if isolated/hardened; high if internet-exposed or logs leak prompts. Low vendor lock-in at model/API layer; hardware/cloud lock-in possible. Secret-sensitive private inference, high-volume background tasks, self-hosted code/RAG models. Tiny VPS without GPU; tasks needing proprietary frontier accuracy. Use OpenAI-compatible endpoint behind internal gateway; disable prompt logs or protect them; patch promptly. vLLM, GitHub, docs mirror — checked 2026-08-07.
Replicate Hosted public/private ML models, image/video/audio/LLM predictions, webhooks. Broad model marketplace, quick prototypes, official models with stable API. Per-second compute can be costly; public model supply-chain and license quality varies. Pay-as-you-go compute time; model pages show estimate; public models billed by run time/hardware or model-specific units. Great for media/prototypes; LLM quality depends hosted model and cold starts. Model licenses/copyright vary; user responsible for generated media/code use. Check retention/privacy for predictions/files; not secret-safe by default. Medium-high for uploaded media/private data. Platform/model marketplace lock-in; export possible for open models. Vision/media experiments, non-secret model prototypes. Secret docs, production LLM substrate at scale without cost model. Prefer official models for stable API; set webhooks carefully; avoid private keys in inputs. pricing, billing, official models — checked 2026-08-07.
RunPod GPU pods, serverless GPUs, clusters for self-hosted models/vLLM/Ollama. Flexible GPU cost, serverless scale-to-zero, good for self-hosting and batch jobs. Ops burden, cold starts, container security, spot/preemption risk, storage/network extra costs. Official pricing per GPU/hour or serverless per second; serverless bills from worker start until fully stopped; pricing changed June 2026 with Flex/Active workers. Quality depends served model; latency depends worker/cold start/GPU. Self-hosted model licenses and local policy enforcement. Data stays in chosen deployment/provider infra; still third-party GPU cloud, so encrypt and avoid raw secrets unless contract. Medium; lower with encrypted volumes/private network; high on community/spot for secrets. GPU provider lock-in moderate; Docker/vLLM reduces. Self-hosted open models, experiments, batch inference, private-ish code/RAG when contract acceptable. Highest-security secrets without dedicated secure cloud/contract. Harden images, no secrets baked into containers, restrict network, clean volumes. pricing, serverless pricing, pricing update — checked 2026-08-07.
AWS Bedrock / Azure AI Foundry / cloud marketplaces Managed access to Anthropic, Meta, Mistral, Amazon, OpenAI and other models with IAM/private networking/regional controls. Enterprise procurement, IAM, VPC/PrivateLink, audit, compliance, data residency options. Higher complexity/cost; model availability by region; service-tier lock-in. Bedrock on-demand per token, batch 50% lower for select FMs, provisioned/reserved tiers; exact per-model rates region-dependent. Quality depends model; latency and quotas by region/tier. Underlying model + cloud provider policies both apply. Stronger enterprise controls possible; configure logs carefully. Low-medium with private networking/ZDR/contract; medium if default logging broad. Strong cloud lock-in; better geography controls. High-security enterprise route when Synapolis standardizes on cloud IAM/contracts. Casual cheap experiments if direct API cheaper and data non-sensitive. Use when compliance/procurement matters more than cheapest token. Bedrock pricing, model availability — checked 2026-08-07.

Security notes for agent substrate[edit | edit source]

  • Secrets in prompts are exfiltration, not "context". Treat API keys, seed phrases, bearer tokens, SSH keys, cookies, private repo credentials, customer data and unpublished business plans as out-of-scope for third-party API calls unless the contour explicitly allows it.
  • BYOK aggregators reduce integration work but increase blast radius: the aggregator can see prompts and may hold provider keys, route metadata, spend history and error traces. Use direct provider APIs or self-hosted gateway for sensitive work.
  • Browser/search tools are prompt-injection surfaces. A web page can instruct the model to reveal memory, call tools, change code, or ignore policy. Search/extract output must be treated as untrusted data, not as system instructions.
  • RAG vectors are derived sensitive data. Embeddings can leak semantic content and should be partitioned by sensitivity, model, tenant and deletion policy.
  • Low-policy/open-weight does not remove law, copyright, privacy, export-control or platform abuse obligations. It only shifts enforcement from provider to Synapolis.
  • For high-security agents: prefer local/self-hosted inference, no prompt logging, encrypted volumes, network egress allowlists, per-agent virtual keys, spend limits, and separate browser/search sandbox from credentialed tools.

Practical routing policy[edit | edit source]

Workload Default model lane Cheap lane Secure lane Search/RAG lane
Simple background classify/summarize Gemini Flash / OpenAI mini / Mistral small DeepSeek Flash, Qwen Flash, Groq 8B/70B Ollama/vLLM small model Brave/Tavily only if current facts needed
Code patch/review Claude Sonnet/Opus or OpenAI strong code/reasoning Qwen/Kimi/DeepSeek via Fireworks/Together/Groq after tests Local vLLM code model for private repo pre-pass; human-gated external call if needed Repo-local grep/tests first; web search only for public docs
Long analysis/writing OpenAI/Anthropic/Gemini Pro/xAI DeepSeek/Qwen/Mistral Local model for confidential drafts, then redacted external polish Perplexity/Brave/Tavily with citations
Current research Perplexity Sonar + verification LLM Brave/Tavily + cheap summarizer Self-hosted reader over approved sources Voyage/Jina/Cohere rerank for local corpus
Secret-sensitive operation No third-party by default No cheap external lane Local/self-hosted or enterprise ZDR/contract only Internal docs only; no open web mixed with credentials

Sources checked[edit | edit source]

Primary official docs/pricing/policy pages were preferred. Non-official pages were used only where official pages are dynamic, model-card-specific, or sparse, and the table marks those as orientation rather than contract source.