Jump to content
Main menu
Main menu
move to sidebar
hide
Navigation
Main page
Recent changes
Random page
Help about MediaWiki
wikibase
Search
Search
English
Create account
Log in
Personal tools
Create account
Log in
Pages for logged out editors
learn more
Contributions
Talk
Editing
Synapolis/Agent Substrate API Catalog
(section)
Page
Discussion
English
Read
Edit
Edit source
View history
Tools
Tools
move to sidebar
hide
Actions
Read
Edit
Edit source
View history
General
What links here
Related changes
Special pages
Page information
Warning:
You are not logged in. Your IP address will be publicly visible if you make any edits. If you
log in
or
create an account
, your edits will be attributed to your username, along with other benefits.
Anti-spam check. Do
not
fill this in!
== Сравнительная таблица == Сокращения: ''ZDR'' — zero data retention; ''BYOK'' — bring your own key; ''MTok'' — 1 млн токенов; ''официально'' означает, что строка опирается на публичные docs/pricing/policy provider-а, но не заменяет contract review. {| class="wikitable sortable" ! API/provider ! Типы возможностей ! Сильные стороны ! Слабые стороны ! Цена/стоимость ориентир и единица ! Качество / latency / context / tool-use / code ! Policy / copyright / content limits ! Data retention / training / privacy controls ! Secret leakage risk ! Lock-in / availability / geography ! Recommended use in Synapolis agents ! Not recommended for ! Integration notes ! Source links / date checked |- | OpenAI | LLM/chat/responses, reasoning, tool-use, code, embeddings, speech/vision/image, batch, regional processing. | Сильная универсальность, mature API/SDK, хорошее tool/function calling, embeddings дешёвые, enterprise privacy/ZDR options. | Premium frontier costs; policy enforcement; модельный роутинг и названия быстро меняются; некоторые endpoints/features могут иметь особые retention limits. | От ~$0.05/MTok input на nano/cheap tiers до premium frontier; embeddings text-embedding-3-small около $0.02/MTok, 3-large около $0.13/MTok; batch обычно дешевле. Проверять live pricing. | Высокое качество для agents/code/reasoning; latency зависит от model/service tier; context у новых семейств large. | Строгие universal usage policies; ownership terms assign output to user where law allows; copyright/IP risk всё равно остаётся у интегратора по downstream use. | API data не используется для training без opt-in; default retention до 30 дней для abuse/service, ZDR для qualifying org/endpoints; regional processing с uplift для eligible models. | Средний: прямой API лучше BYOK-агрегатора, но secrets всё равно нельзя слать; включать redaction и project-level keys. | US provider; Azure/OpenAI/AWS Bedrock/regional options уменьшают, но не убирают lock-in. | Default high-quality substrate; code review, synthesis, structured tool-use, embeddings baseline. | Неподготовленные secret-bearing prompts; use cases без human oversight в regulated/high-stakes областях. | Использовать Responses API/tool schemas, batch for background, cache, service tiers; отдельные org/project keys. | [https://developers.openai.com/api/docs/pricing pricing], [https://developers.openai.com/api/docs/guides/your-data data controls], [https://openai.com/policies/usage-policies/ usage policies], [https://openai.com/policies/row-terms-of-use/ terms] — checked 2026-08-07. |- | Anthropic Claude | LLM/chat/messages, extended thinking, tool-use, computer/code workflows, managed agents/Claude Code surfaces. | Очень сильный coding/repo reasoning, long-form analysis, cautious tool-use, enterprise posture. | Дороже low-cost конкурентов; policy stricter; model availability/rate limits can bottleneck; consumer Claude Code privacy differs from API/enterprise. | Claude pricing page: Sonnet 5 introductory ~$2/$10 per MTok input/output through 2026-08-31 then ~$3/$15; Haiku lower, Opus/Fable higher. Проверять live. | High quality for code agents and analysis; context/model family varies; latency often higher than small models. | Strong AUP/safety; явные disallowed/high-risk use constraints; copyright litigation/news risk exists, but API output/IP terms need contract review. | API commercial products: not used for training by default; standard deletion within 30 days except files/violations/law/agreements; ZDR available for eligible API/enterprise. | Средний-низкий при API+ZDR; высокий для consumer/free/chat or prompts containing repo secrets. | US provider, strong AWS/GCP/Bedrock routes; potential lock-in through Claude Code workflows. | Primary substrate for serious coding agents, careful writing/analysis, high-risk tool plans with review. | Cheap background classification; low-policy use; secret-bearing workflows without enterprise/ZDR. | Prefer direct API/Bedrock/Vertex where contracts fit; use prompt caching/batch; tool boundaries explicit. | [https://platform.claude.com/docs/en/about-claude/pricing pricing], [https://platform.claude.com/docs/en/about-claude/models/overview models], [https://platform.claude.com/docs/en/manage-claude/api-and-data-retention retention], [https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training training], [https://www.anthropic.com/news/usage-policy-update usage policy] — checked 2026-08-07. |- | Google Gemini / Vertex AI / Agent Platform | Gemini API, Vertex AI, multimodal text/image/video/audio, grounding with Google Search, agents, embeddings, speech/vision via Google Cloud. | Strong multimodal, large context, cheap Flash tiers, Google Cloud IAM/VPC/region/data tooling, integrated Search grounding. | Policy/terms differ between free AI Studio, paid Gemini API, Vertex/Agent Platform; store=true defaults and logging settings need care. | Gemini API snippets show Flash-like paid tiers around $0.15/$1.25 MTok and higher Pro tiers; Vertex/Agent Platform Gemini 3.x Pro/Flash pricing can be $1.5-$4 input and $9-$18 output depending model/context/region. Search grounding after free quota about $14/1K requests. | Very good latency/cost on Flash; Pro stronger for analysis; excellent multimodal; tool-use acceptable but integration choices matter. | Google generative AI prohibited use policy; stricter around illegal/harmful, privacy, deception, explicit sexual content depending product. | Paid Gemini API: no training on prompts/responses per ZDR docs; ZDR requires avoiding store/session/logging features; Agent Platform request-response logging disabled by default but Interactions API may default store=true unless set false. | Средний: Google Cloud controls can be strong, but accidental storage/logging is easy. | Strong lock-in if using Vertex/Agent Platform/Search grounding; good regional/cloud availability. | Cheap massive background agents, multimodal ingestion, grounded search, GCP-native secure deployments. | Low-policy needs; workflows that cannot tolerate Google account/project policy changes. | Explicitly set store=false where applicable; separate AI Studio/free from paid API; use IAM/service accounts. | [https://ai.google.dev/gemini-api/docs/pricing Gemini API pricing], [https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing Agent Platform pricing], [https://ai.google.dev/gemini-api/docs/zdr Gemini ZDR], [https://docs.cloud.google.com/gemini-enterprise-agent-platform/resources/zero-data-retention Agent Platform ZDR], [https://policies.google.com/terms/generative-ai/use-policy policy] — checked 2026-08-07. |- | xAI Grok | LLM/chat/code, reasoning, tool-use, image/video/voice APIs. | Grok 4.5 marketed for code, non-hallucination and agentic tool calling; comparatively lenient brand positioning; OpenAI-compatible style. | Smaller enterprise/privacy track record than OpenAI/Google/Anthropic; policy and terms changed recently; limited independent operational history. | Official pricing: grok-4.5 short context about $2 input / $6 output MTok, long context $4/$12; grok-build lower; media priced per image/sec. | Potentially strong code/agentic tasks; 500K+ context on flagship; latency depends on capacity. | AUP says maximize user control while requiring legal/responsible/safe use; not "uncensored" for illegal/harmful use. | Privacy policy applies; enterprise terms for API/business; verify API retention/training in contract/docs before secrets. | Средний-высокий until org-specific privacy terms are confirmed. | US provider; model catalog narrower; availability/capacity may fluctuate. | Experimental coding agents, writing/analysis alternative, low-friction public-data tasks. | Secret-sensitive or regulated workflows without enterprise terms; sole provider dependency. | Use as routed option behind gateway with fallback; log model/version and policy date. | [https://docs.x.ai/developers/pricing pricing], [https://docs.x.ai/developers/models models], [https://x.ai/legal/acceptable-use-policy AUP], [https://x.ai/legal/terms-of-service-enterprise enterprise terms], [https://x.ai/legal/privacy-policy privacy] — checked 2026-08-07. |- | Mistral AI | LLM/chat, agents Studio/Vibe, code Vibe, embeddings, OCR, fine-tuning, EU enterprise. | EU provider, open/proprietary mix, good price-performance, privacy controls, self-hostable/open-weight ecosystem. | Public pricing page may require docs/console for detailed model rows; ZDR only on Scale/stateless API; some products stateful. | API starts low; exact model prices vary. Official docs/pricing should be checked per model; third-party summaries place cheap tiers around $0.10+/MTok and frontier higher. | Good multilingual/code/general; latency good on smaller models; tool-use maturing. | Mistral usage/legal docs; generally less restrictive than Anthropic/OpenAI but not law-free. | API privacy controls for training/retention; ZDR available only Scale and stateless calls, not Vibe/agents/conversations/batch/stateful products. | Средний-низкий for Scale ZDR stateless; средний for normal API; high for stateful agent products with secrets. | EU geography advantage; lock-in moderate if using Studio/Vibe agents. | EU-friendly general agents, cheaper writing, open-weight/local transition path. | High-security unless ZDR/contract in place; tasks requiring best-in-class coding certainty. | Prefer stateless API for private use; document if using Labs/free tiers. | [https://mistral.ai/pricing/ pricing], [https://docs.mistral.ai/models/overview models], [https://docs.mistral.ai/admin/monitor-comply/privacy-data-controls privacy controls], [https://help.mistral.ai/en/articles/347612-can-i-activate-zero-data-retention-zdr ZDR], [https://legal.mistral.ai/terms terms] — checked 2026-08-07. |- | Cohere | Command LLMs, enterprise RAG, Embed, Rerank, Model Vault/dedicated deployments. | Strong RAG/rerank/multilingual enterprise story; dedicated Model Vault for secure predictable inference. | Not usually first pick for frontier coding; pricing page mixes API and dedicated infra; generation catalog less buzz than frontier labs. | Rerank 4 Pro via OpenRouter listed ~$0.0025/search; dedicated Model Vault examples $4-$10/hour instances; older Embed/Rerank per-token varies by platform. | Excellent retrieval/reranking; Command useful for enterprise RAG; code weaker than Claude/OpenAI top tiers. | Enterprise acceptable-use terms; not a low-policy provider. | Enterprise privacy/security controls; verify per API/platform and cloud marketplace route. | Low-medium for dedicated/vault; medium for API; never send raw secrets into RAG rerank unless approved. | Canada/enterprise/cloud marketplace availability; some lock-in in RAG pipeline tuning. | RAG rerank, multilingual embeddings, private enterprise retrieval substrate. | Primary code agent model; cheap creative chat. | Use Cohere Rerank after cheap vector retrieval; cap docs/query and log token/search cost. | [https://cohere.com/pricing pricing], [https://docs.cohere.com/docs/how-does-cohere-pricing-work pricing docs], [https://docs.cohere.com/docs/models models], [https://openrouter.ai/cohere/rerank-4-pro rerank price reference] — checked 2026-08-07. |- | DeepSeek | LLM/chat/reasoning/code, OpenAI-compatible endpoint, long context, low-cost models. | Extremely strong price-performance for coding/reasoning; cheap background agents; large context on newer models. | Chinese jurisdiction/censorship/security concerns; pricing page warns of future increases; availability/payment/region risk. | Official docs: V4 Flash about $0.14 input cache miss / $0.28 output MTok; V4 Pro about $0.435/$0.87 MTok with very cheap cache-hit input; price increase warning. | Good code/reasoning for cost; latency/capacity can vary; tool-use depends wrapper. | Content/policy/censorship profile differs from US/EU providers; avoid political/legal-sensitive assumptions. | Privacy policy exists; do not assume Western enterprise ZDR. Treat as non-secret public-data substrate unless contract says otherwise. | Высокий for secrets/IP-sensitive prompts; acceptable for public/low-risk tasks after redaction. | China-linked; sanctions/export/procurement/geography risk; strong lock-in if prompts tuned to DeepSeek quirks. | Cheap background, code draft, long-context public analysis, fallback in router. | Secrets, customer data, regulated workloads, geopolitical/policy-sensitive analysis without cross-check. | Use via LiteLLM/OpenAI-compatible route with strict redaction; verify model IDs and peak pricing. | [https://api-docs.deepseek.com/quick_start/pricing/ pricing], [https://cdn.deepseek.com/policies/en-US/deepseek-privacy-policy.html privacy] — checked 2026-08-07. |- | Qwen / Alibaba Cloud Model Studio | Qwen text/code/multimodal/audio/video via Alibaba Cloud Model Studio; open-weight Qwen locally or via providers. | Strong open-weight ecosystem, cheap and capable multilingual/code models, cloud + self-host path. | International availability/region/payment and policy/censorship concerns; docs/pricing by region; China-linked procurement risk. | Official Model Studio pay-as-you-go by model; third-party summaries cite Qwen Flash around $0.10/$0.40 MTok, Plus around $0.40/$2.40, 397B higher. Verify live Alibaba Cloud region. | Very good price-performance; Qwen Coder strong; latency depends host. | Alibaba Cloud policies; Chinese jurisdiction/censorship risk; not for illegal/harmful use. | Verify Alibaba Cloud data handling per account/region; open weights self-host avoid API retention but shift ops risk to Synapolis. | High for direct cloud with secrets unless contract; low if self-hosted and isolated. | China/Singapore/international region constraints; open-weight reduces lock-in. | Cheap code/background via open weights; self-host candidates for private tasks. | Sensitive Western customer data/direct secrets on Alibaba-hosted API. | Prefer self-host/open-weight when privacy matters; use cloud for public/low-risk throughput. | [https://www.alibabacloud.com/help/en/model-studio/models models], [https://www.alibabacloud.com/help/en/model-studio/model-pricing pricing], [https://modelstudio.alibabacloud.com/ Model Studio] — checked 2026-08-07. |- | Moonshot / Kimi | Kimi LLM/code/long-context, Kimi K-series; API and open-weight/aggregator routes. | Strong code/agentic benchmarks and long-context claims; open-weight Kimi variants broaden hosting options. | Direct official pricing/resources can be less enterprise-transparent; China-linked risk; top models may be expensive. | Kimi resources page: K3 about $3 input cache miss / $0.30 cache hit / $15 output MTok; K2.6 about $0.95 miss / $0.16 hit / $4 output; OpenRouter K3 about $2.50/$14. | Strong coding and long-horizon agentic workflows if model claims hold; latency depends host/provider. | Treat as low-policy only in the sense of less Western SaaS friction; still legal/policy and censorship risks. | Do not assume ZDR on direct API/aggregators; self-hosted open weights give better control. | High for secrets unless self-hosted and audited. | China-linked and aggregator-dependent availability; open weights reduce vendor lock-in. | Coding experiments, public repo agents, cost/performance benchmark lane. | Secret-sensitive, regulated, political/censorship-critical outputs without independent verification. | Prefer OpenRouter/Together/Fireworks only for non-secret eval; pin provider route for reproducibility. | [https://www.moonshot.ai/ Moonshot], [https://www.kimi.com/resources/kimi-k3-pricing Kimi K3 pricing], [https://www.kimi.com/resources/kimi-k2-6-pricing Kimi K2.6 pricing], [https://openrouter.ai/moonshotai/kimi-k3 OpenRouter K3] — checked 2026-08-07. |- | Together AI | Serverless inference, dedicated endpoints, GPU clusters, fine-tuning, embeddings/rerank/moderation, open models. | Broad open-model catalog, serverless-to-dedicated path, privacy settings/ZDR, good for routing experiments. | Cost varies by model; passthrough third-party model privacy can differ; GPU/dedicated economics require capacity planning. | Serverless per MTok by model; public summaries show ~$0.03-$7+/MTok; GPU clusters/dedicated by GPU-hour; check live pricing. | Good latency/throughput; quality follows selected model; supports many code/open models. | Provider policy plus third-party/open-model licenses; not a single uniform content policy. | Privacy settings can disable storage/training and enable ZDR; docs note third-party model providers/passthrough need attention. | Medium; lower with ZDR and direct provider route, higher with passthrough/unknown models. | US provider; open catalog reduces model lock-in, platform lock-in remains. | Open-model benchmark lane, cheap background, dedicated endpoints for stable workloads. | Secrets unless ZDR/contract and no passthrough ambiguity. | Use model allowlists, privacy settings, spend caps; record exact model/provider. | [https://www.together.ai/pricing pricing], [https://docs.together.ai/docs/privacy-and-security privacy/security], [https://www.together.ai/terms-of-service terms/ZDR], [https://www.together.ai/privacy privacy] — checked 2026-08-07. |- | Fireworks AI | Serverless/on-demand/reserved inference, open models, embeddings, fine-tuning, OpenAI/Anthropic-compatible APIs. | High-throughput open-model serving, ZDR posture for open models, good enterprise options, batch 50% pricing. | Exact prices per model/tier; prepaid billing; hosted model quality follows open model. | Official: serverless per MTok; embeddings as low as $0.008-$0.1/MTok by model/size; batch 50% of serverless; H100/on-demand pricing separate. | Good latency/throughput for open models; code quality depends Qwen/Kimi/DeepSeek/Llama selected. | Open-model licenses and Fireworks terms; not an illegal/uncensored bypass. | Privacy policy says no AI training without opt-in; docs describe ZDR for most services, Response API has retention caveat. | Medium-low for open-model ZDR; medium if response/conversation state is used. | US provider; platform lock-in moderate; self-host transition possible by using open weights. | Cheap public-data agents, code model hosting, batch inference, fine-tuned open models. | Highest security unless contract; workflows needing proprietary frontier quality. | Use standard/priority/fast tier explicitly; batch background jobs; avoid stateful retention for secrets. | [https://fireworks.ai/pricing pricing], [https://docs.fireworks.ai/serverless/pricing serverless pricing], [https://docs.fireworks.ai/guides/security_compliance/data_handling ZDR/data handling], [https://fireworks.ai/privacy-policy privacy] — checked 2026-08-07. |- | GroqCloud | Fast inference for selected open models, Whisper/STT, TTS, tool-use if model supports. | Very low latency/tokens-per-second, cheap small/medium open models, self-serve ZDR controls. | Narrower model catalog; not always best quality; rate limits/capacity can shape production behavior. | Official docs: Llama 3.1 8B ~$0.05/$0.08 MTok; Llama 3.3 70B ~$0.59/$0.79; GPT-OSS 120B ~$0.15/$0.60; Whisper by audio hour. | Excellent latency; quality follows model; context often 131K on listed models. | Acceptable-use/responsible AI policies; open-model licenses apply. | Docs: usage metadata collected; customer data not retained by default except opt-in/persistence/reliability; ZDR self-serve. | Medium-low with ZDR; still third-party cloud, no raw secrets. | US provider, specialized hardware; model choice/capacity lock-in. | Fast cheap background agents, streaming UX, STT transcription, low-latency classifiers. | Hard reasoning/code requiring frontier accuracy; secret workflows without isolation. | Use as latency tier behind router; fallback on quality-sensitive tasks. | [https://groq.com/pricing pricing], [https://console.groq.com/docs/models models], [https://console.groq.com/docs/your-data data], [https://console.groq.com/docs/legal legal] — checked 2026-08-07. |- | Cerebras | Ultra-fast inference for selected open/frontier-ish models; developer and paid tiers. | Very high speed claims; useful for latency-sensitive agents and batch acceleration. | Smaller catalog; detailed per-model token rates may require console; enterprise maturity narrower than hyperscalers. | Official pricing page: free trial $5 credits; developer self-serve starts around $10; public docs emphasize flexible pricing. Historical public rates were ~$0.10/MTok 8B and ~$0.60/MTok 70B, verify live. | Very low latency/high throughput; quality follows model selection. | Standard AUP/contracts; open-model licenses. | Verify retention/training terms in contract/trust center before secrets. | Medium until ZDR/contract confirmed. | US provider, specialized infra; availability/catalog lock-in. | Latency-sensitive public-data agents, fast eval loops. | Secret/high-compliance use without terms; broad multimodal tasks. | Treat as speed lane in router; benchmark live latency with Synapolis prompts. | [https://www.cerebras.ai/pricing pricing], [https://www.cerebras.ai/ press/site], [https://www.cerebras.ai/press-release/cerebras-launches-the-worlds-fastest-ai-inference historical pricing context] — checked 2026-08-07. |- | Perplexity Sonar | Search-grounded chat/completions, citations, answer engine, OpenAI-compatible API. | Useful search/browser substrate with citations and low integration cost. | More expensive than raw search; citation quality must be verified; not a browser automation replacement. | Official docs: Gateway billed per token at model rates; Sonar pay-as-you-go no subscription. Third-party summaries cite base ~$1/$1 MTok plus request/search fees; verify live model card. | Good for current factual answers; latency includes search; not best for code edits. | Must respect source copyright/ToS; not a license to copy pages. | Verify API data terms; search queries may reveal intent/secrets. | Medium-high: search prompts often contain sensitive business questions. | US provider; source index/model routing dependency. | Current research with citations, pre-RAG discovery, fact-check lane. | Secret/internal search queries; legal/medical/financial final answers without human/source review. | Always store cited URLs/date; use as input to human/agent verification, not final truth. | [https://docs.perplexity.ai/docs/getting-started/pricing pricing], [https://docs.perplexity.ai/docs/sonar/quickstart Sonar quickstart] — checked 2026-08-07. |- | Tavily | Web search, extract, crawl/research endpoints for AI agents/RAG. | Agent-oriented search/extraction API; simple credit pricing; good for web research pipelines. | Credits can burn quickly on /research; source quality/copyright/prompt injection risks remain. | Free 1,000 credits/month; PAYG ~$0.008/credit; monthly ~$0.0075-$0.005/credit; basic/advanced/research consume different credits. | Latency/search quality good for agent use; not an LLM by itself. | Web content license/copyright responsibility remains with integrator. | Search queries/extracted pages sent to Tavily; verify privacy terms for sensitive topics. | Medium: queries can leak plans; never include credentials. | SaaS provider; index/extraction behavior lock-in moderate. | Search/RAG substrate, source discovery, monitoring. | Secret-bearing browsing; executing page instructions. | Sanitize queries; isolate browser/search tool from credentialed tools; cite sources. | [https://www.tavily.com/pricing pricing], [https://docs.tavily.com/documentation/api-credits credits], [https://tavily.com/ product] — checked 2026-08-07. |- | Brave Search API | Independent web/news/images/LLM context/search API. | Own index, predictable pricing, citations/LLM context, privacy-friendly brand. | Not Google SERP; niche/technical coverage may differ; lower QPS on small plans. | Official: Search $5 per 1,000 requests, includes $5 monthly credits; enterprise custom. | Good raw search/grounding; latency and completeness vary by query. | Search result snippets/URLs only; downstream copyright/ToS still apply. | Queries go to Brave; privacy better than many search APIs but not secret-safe. | Medium-low for non-secret queries; medium for strategy queries. | Vendor/index lock-in; global web availability. | Default cheap independent search for agents and RAG. | Need exact Google SERP/ads/local fidelity; secrets in search terms. | Use LLM Context endpoint where token-efficient; keep citation URLs. | [https://brave.com/search/api/ pricing/API], [https://brave.com/learn/best-search-api-2026/ pricing explanation] — checked 2026-08-07. |- | SerpAPI | Google/Bing/Yahoo/Maps/Shopping/etc SERP JSON. | High SERP fidelity, mature JSON schema, legal shield on paid tiers. | More expensive than Brave/SearchAPI; not AI-specific answer engine. | Starter $25/month for 1,000 searches; Developer $75/5,000; Production $150/15,000; Big Data $275/30,000. | Good SERP extraction; latency tied to live search. | Public search collection/legal shield does not license downstream content use. | Queries expose intent; account logs/retention need vendor review. | Medium. | US provider; Google SERP dependency. | SEO/Google-specific search tasks, monitoring exact SERP features. | Cheap high-volume generic search; secret queries. | Use only when exact SERP matters; otherwise Brave/Tavily. | [https://serpapi.com/pricing pricing], [https://serpapi.com/ product/legal shield] — checked 2026-08-07. |- | SearchAPI.io | Google SERP and related search APIs. | Cheaper Google SERP alternative with SLA/legal protection on higher plans. | Smaller ecosystem than SerpAPI; plan semantics and enhanced speed need modeling. | Developer $40/month, about $4/1K; Production $100/month, about $3/1K; BigData $250/month, about $2.5/1K; Octo 1M $1.5/1K. | Good for SERP/RAG discovery; verify fields for required verticals. | Same SERP/copyright downstream caveats. | Queries can leak sensitive plans. | Medium. | Vendor + Google SERP dependency. | Cost-sensitive Google SERP substrate. | Non-Google broad web answer generation; secrets. | Benchmark against SerpAPI for needed SERP features before switching. | [https://www.searchapi.io/pricing pricing], [https://www.searchapi.io/ product] — checked 2026-08-07. |- | Voyage AI | Embeddings, multimodal embeddings, rerankers, Batch API. | Excellent retrieval-specialist provider; cheap rerank token pricing and free initial rerank allowance. | Narrow scope; acquired/operated with MongoDB ecosystem; not a chat model. | Rerank-2.5 about $0.05/MTok, lite $0.02/MTok, first 200M rerank tokens free; older embeddings: voyage-3.5 $0.06/MTok, lite $0.02/MTok. | High RAG quality; latency good; context varies by model. | Retrieval content copyright remains integrator responsibility. | Check docs/contract; embeddings encode sensitive content, so treat vectors as derived sensitive data. | Medium: documents/queries sent to third party and vectors may leak semantics. | Vendor lock-in via vector dimensions/model behavior; manageable with abstraction. | RAG indexing/rerank for Synapolis docs, public corpora, non-secret knowledge. | Raw secrets, private keys, personal data without approval. | Store model/version/dimension; separate indexes by sensitivity; rotate if changing model. | [https://docs.voyageai.com/docs/pricing pricing], [https://www.mongodb.com/docs/voyageai/models/ models] — checked 2026-08-07. |- | Jina AI | Embeddings, rerankers, Reader, DeepSearch, multimodal search foundation models. | Strong multilingual/multimodal retrieval, Reader for web-to-markdown/JSON, open research/models. | Pricing details may be dashboard/package-based; less enterprise standard than hyperscalers. | Reranker/Embedding APIs begin with 10M free tokens per new API key; packages beyond. Reader/token pricing varies; verify account pricing. | Good retrieval quality, especially multilingual/multimodal; not primary chat LLM. | Web reader/crawler must respect website ToS/copyright. | API docs need review for retention; embeddings/reader inputs are sensitive derived data. | Medium. | Germany/EU-facing but global SaaS; model/vector lock-in. | RAG, reader/extraction, multilingual retrieval. | Secrets in docs/pages; high-compliance without contract. | Use Reader in isolated pipeline; don't execute extracted instructions. | [https://jina.ai/ product], [https://jina.ai/reranker/ reranker], [https://jina.ai/en-US/embeddings/ embeddings] — checked 2026-08-07. |- | OpenRouter | Aggregated LLM gateway, 400+ models, routing/fallbacks, BYOK/paygo, ZDR filters. | Fast model discovery, unified API, fallback/routing, privacy controls can restrict ZDR providers. | Adds extra trust boundary; provider logging varies; model/provider route can change unless pinned; BYOK concentrates keys in gateway. | Pay-as-you-go per model, claims transparent pricing and many free models; check each model card. | Quality/latency depends provider route; routing can improve availability but reduce reproducibility. | Policies differ by underlying provider; OpenRouter is not a policy bypass. | Docs: prompt retention opt-in; provider logging differs; ZDR can be enforced globally/per model group/guardrail/request. | High if using BYOK with secrets or prompt logging; medium with ZDR/non-secret prompts. | Strong platform lock-in for routing/account credits; reduces model lock-in. | Non-secret evaluation, fallback router, public-data price benchmarking. | High-security/secret-sensitive direct workflows; regulated data unless contract and ZDR routing are audited. | Pin model+provider when reproducibility matters; enable ZDR; disable prompt logging; never route secrets. | [https://openrouter.ai/models models], [https://openrouter.ai/pricing pricing], [https://openrouter.ai/docs/quickstart quickstart], [https://openrouter.ai/docs/guides/privacy/data-collection data], [https://openrouter.ai/docs/guides/privacy/provider-logging provider logging], [https://openrouter.ai/docs/guides/features/zdr ZDR] — checked 2026-08-07. |- | LiteLLM Proxy | Self-hosted LLM gateway/router, virtual keys, spend tracking, budgets, provider abstraction. | Good Synapolis control plane candidate: central keys, budgets, model allowlists, audit, OpenAI-compatible interface. | Security-critical service; past/current vulnerabilities possible; misconfig can expose master keys or allow over-broad routing. | OSS/self-host infra cost; provider token costs pass through; spend tracking uses model cost map. | Latency overhead small but depends routing; quality depends model; supports many providers. | Underlying provider policies apply; LiteLLM itself does not make unsafe use safe. | Privacy depends deployment DB/logging/proxy config and underlying providers. | High if proxy breached; low-medium if self-hosted locked down with per-agent virtual keys and no prompt logging. | Reduces provider lock-in; creates internal gateway lock-in/config dependency. | Synapolis internal model gateway, cost control, per-agent budgets, provider failover. | Running exposed to public internet without auth, patching, DB hardening, secret management. | Put behind private network; rotate master key; virtual keys per agent; redact logs; update pricing map. | [https://docs.litellm.ai/docs/proxy/virtual_keys virtual keys], [https://docs.litellm.ai/docs/proxy/cost_tracking spend tracking], [https://docs.litellm.ai/docs/proxy/users budgets] — checked 2026-08-07. |- | Ollama | Local/cloud open-model runner, embeddings, tool-capable models, offline mode. | Simple local deployment, offline possible, no per-token API cost, data stays local if actually local. | Quality/latency constrained by local hardware; operational burden; cloud features differ from local privacy. | Software free; cost is hardware/electricity/ops. Cloud models/regions may have separate pricing. | Good for small/medium models and private preprocessing; frontier quality usually lower unless large hardware. | Open model licenses vary; self-hosting removes SaaS policy but not law/copyright obligations. | Local/offline: no third-party retention; cloud: check Ollama data handling. | Low if isolated local; medium if exposed API or cloud. | Minimal vendor lock-in; model format/library ecosystem lock-in low. | High-security local summarization/classification, embeddings, private repo pre-pass. | Frontier coding decisions alone; exposed unauthenticated local API. | Bind to localhost/private network; model allowlist; monitor VRAM/RAM; do not mix browser tools with secrets. | [https://ollama.com/ product/data], [https://ollama.com/library model library], [https://github.com/ollama/ollama GitHub] — checked 2026-08-07. |- | vLLM | Self-hosted high-throughput LLM serving, OpenAI-compatible server, batching, prefix caching, quantization/distributed. | Production-grade serving engine for open weights; good throughput; keeps data in Synapolis-controlled infra. | Requires GPU ops, model security, patching, observability; not a model by itself. | OSS; cost is GPU/hour + storage + ops. RunPod/Lambda/CoreWeave/local GPUs determine spend. | Excellent throughput/latency when tuned; quality depends selected model. | Open model licenses and local policy enforcement are Synapolis responsibility. | No third-party API retention if fully self-hosted; logs and traces become local sensitive artifacts. | Low if isolated/hardened; high if internet-exposed or logs leak prompts. | Low vendor lock-in at model/API layer; hardware/cloud lock-in possible. | Secret-sensitive private inference, high-volume background tasks, self-hosted code/RAG models. | Tiny VPS without GPU; tasks needing proprietary frontier accuracy. | Use OpenAI-compatible endpoint behind internal gateway; disable prompt logs or protect them; patch promptly. | [https://vllm.ai/ vLLM], [https://github.com/vllm-project/vllm GitHub], [https://nm-vllm.readthedocs.io/ docs mirror] — checked 2026-08-07. |- | Replicate | Hosted public/private ML models, image/video/audio/LLM predictions, webhooks. | Broad model marketplace, quick prototypes, official models with stable API. | Per-second compute can be costly; public model supply-chain and license quality varies. | Pay-as-you-go compute time; model pages show estimate; public models billed by run time/hardware or model-specific units. | Great for media/prototypes; LLM quality depends hosted model and cold starts. | Model licenses/copyright vary; user responsible for generated media/code use. | Check retention/privacy for predictions/files; not secret-safe by default. | Medium-high for uploaded media/private data. | Platform/model marketplace lock-in; export possible for open models. | Vision/media experiments, non-secret model prototypes. | Secret docs, production LLM substrate at scale without cost model. | Prefer official models for stable API; set webhooks carefully; avoid private keys in inputs. | [https://replicate.com/pricing pricing], [https://replicate.com/docs/topics/billing billing], [https://replicate.com/docs/topics/models/official-models official models] — checked 2026-08-07. |- | RunPod | GPU pods, serverless GPUs, clusters for self-hosted models/vLLM/Ollama. | Flexible GPU cost, serverless scale-to-zero, good for self-hosting and batch jobs. | Ops burden, cold starts, container security, spot/preemption risk, storage/network extra costs. | Official pricing per GPU/hour or serverless per second; serverless bills from worker start until fully stopped; pricing changed June 2026 with Flex/Active workers. | Quality depends served model; latency depends worker/cold start/GPU. | Self-hosted model licenses and local policy enforcement. | Data stays in chosen deployment/provider infra; still third-party GPU cloud, so encrypt and avoid raw secrets unless contract. | Medium; lower with encrypted volumes/private network; high on community/spot for secrets. | GPU provider lock-in moderate; Docker/vLLM reduces. | Self-hosted open models, experiments, batch inference, private-ish code/RAG when contract acceptable. | Highest-security secrets without dedicated secure cloud/contract. | Harden images, no secrets baked into containers, restrict network, clean volumes. | [https://www.runpod.io/pricing pricing], [https://docs.runpod.io/serverless/pricing serverless pricing], [https://www.runpod.io/blog/serverless-pricing-update pricing update] — checked 2026-08-07. |- | AWS Bedrock / Azure AI Foundry / cloud marketplaces | Managed access to Anthropic, Meta, Mistral, Amazon, OpenAI and other models with IAM/private networking/regional controls. | Enterprise procurement, IAM, VPC/PrivateLink, audit, compliance, data residency options. | Higher complexity/cost; model availability by region; service-tier lock-in. | Bedrock on-demand per token, batch 50% lower for select FMs, provisioned/reserved tiers; exact per-model rates region-dependent. | Quality depends model; latency and quotas by region/tier. | Underlying model + cloud provider policies both apply. | Stronger enterprise controls possible; configure logs carefully. | Low-medium with private networking/ZDR/contract; medium if default logging broad. | Strong cloud lock-in; better geography controls. | High-security enterprise route when Synapolis standardizes on cloud IAM/contracts. | Casual cheap experiments if direct API cheaper and data non-sensitive. | Use when compliance/procurement matters more than cheapest token. | [https://aws.amazon.com/bedrock/pricing/ Bedrock pricing], [https://docs.aws.amazon.com/bedrock/latest/userguide/models.html model availability] — checked 2026-08-07. |}
Summary:
Please note that all contributions to wikibase may be edited, altered, or removed by other contributors. If you do not want your writing to be edited mercilessly, then do not submit it here.
You are also promising us that you wrote this yourself, or copied it from a public domain or similar free resource (see
Wikibase:Copyrights
for details).
Do not submit copyrighted work without permission!
Cancel
Editing help
(opens in new window)
Toggle limited content width