CLIProxyAPI

mirror of https://github.com/router-for-me/CLIProxyAPI.git synced 2026-02-03 04:50:52 +08:00

Author	SHA1	Message	Date
Ben Vargas	70ee4e0aa0	chore: remove unused httpx sdk package	2025-11-19 21:17:52 -07:00
Ben Vargas	03334f8bb4	chore: revert gitignore change	2025-11-19 20:42:23 -07:00
Ben Vargas	5a2bebccfa	fix: remove duplicate CountTokens stub	2025-11-19 20:00:39 -07:00
Luis Pater	0586da9c2b	refactor(registry): move Gemini 3 Pro Preview model definition to base set	2025-11-20 10:51:16 +08:00
Ben Vargas	3d8d02bfc3	Fix amp v1beta1 routing and gemini retry config	2025-11-19 19:11:35 -07:00
Ben Vargas	7ae00320dc	fix(amp): enable OAuth fallback for Gemini v1beta1 routes AMP CLI sends Gemini requests to non-standard paths that were being directly proxied to ampcode.com without checking for local OAuth. This fix adds: - GeminiBridge handler to transform AMP CLI paths to standard format - Enhanced model extraction from AMP's /publishers/google/models/* paths - FallbackHandler wrapper to check for local OAuth before proxying Flow: - If user has local Google OAuth → use it (free tier) - If no local OAuth → fallback to ampcode.com (charges credits) Fixes issue where gemini-3-pro-preview requests always charged AMP credits even when user had valid Google Cloud OAuth configured.	2025-11-19 18:23:17 -07:00
Ben Vargas	1fb96f5379	docs: reposition Amp CLI as integrated feature for upstream PR - Update README.md to present Amp CLI as standard feature, not fork differentiator - Remove USING_WITH_FACTORY_AND_AMP.md (fork-specific, Factory docs live upstream) - Add comprehensive docs/amp-cli-integration.md with setup, config, troubleshooting - Eliminate fork justification messaging throughout documentation - Prepare Amp CLI integration for upstream merge consideration This positions Amp CLI support as a natural extension of CLIProxyAPI's multi-client architecture rather than a fork-specific feature.	2025-11-19 18:23:17 -07:00
Ben Vargas	897d108e4c	docs: update Factory config with GPT-5.1 models and explicit reasoning levels - Replace deprecated GPT-5 and GPT-5-Codex with GPT-5.1 family - Add explicit reasoning effort levels (low/medium/high) - Remove duplicate base models (use medium as default) - GPT-5.1 Codex Mini supports medium/high only (per OpenAI docs) - Remove older Claude Sonnet 4 (keep 4.5) - Final config: 11 models (3 Claude + 8 GPT-5.1 variants)	2025-11-19 18:23:17 -07:00
Ben Vargas	72d82268e5	fix(amp): filter context-1m beta header for local OAuth providers Amp CLI sends 'context-1m-2025-08-07' in Anthropic-Beta header which requires a special 1M context window subscription. After upstream rebase to v6.3.7 (commit `38cfbac`), CLIProxyAPI now respects client-provided Anthropic-Beta headers instead of always using defaults. When users configure local OAuth providers (Claude, etc), requests bypass the ampcode.com proxy and use their own API subscriptions. These personal subscriptions typically don't include the 1M context beta feature, causing 'long context beta not available' errors. Changes: - Add filterBetaFeatures() helper to strip specific beta features - Filter context-1m-2025-08-07 in fallback handler when using local providers - Preserve full headers when proxying to ampcode.com (paid users get all features) - Add 7 test cases covering all edge cases This fix is isolated to the Amp module and only affects the local provider path. Users proxying through ampcode.com are unaffected and receive full 1M context support as part of their paid service.	2025-11-19 18:23:17 -07:00
Ben Vargas	8193392bfe	Add AMP fallback proxy and shared Gemini normalization - add fallback handler that forwards Amp provider requests to ampcode.com when the provider isn’t configured locally - wrap AMP provider routes with the fallback so requests always have a handler - share Gemini thinking model normalization helper between core handlers and AMP fallback	2025-11-19 18:23:17 -07:00
Ben Vargas	9ad0f3f91e	feat: Add Amp CLI integration with comprehensive documentation Add full Amp CLI support to enable routing AI model requests through the proxy while maintaining Amp-specific features like thread management, user info, and telemetry. Includes complete documentation and pull bot configuration. Features: - Modular architecture with RouteModule interface for clean integration - Reverse proxy for Amp management routes (thread/user/meta/ads/telemetry) - Provider-specific route aliases (/api/provider/{provider}/*) - Secret management with precedence: config > env > file - 5-minute secret caching to reduce file I/O - Automatic gzip decompression for responses - Proper connection cleanup to prevent leaks - Localhost-only restriction for management routes (configurable) - CORS protection for management endpoints Documentation: - Complete setup guide (USING_WITH_FACTORY_AND_AMP.md) - OAuth setup for OpenAI (ChatGPT Plus/Pro) and Anthropic (Claude Pro/Max) - Factory CLI config examples with all model variants - Amp CLI/IDE configuration examples - tmux setup for remote server deployment - Screenshots and diagrams Configuration: - Pull bot disabled for this repo (manual rebase workflow) - Config fields: AmpUpstreamURL, AmpUpstreamAPIKey, AmpRestrictManagementToLocalhost - Compatible with upstream DisableCooling and other features Technical details: - internal/api/modules/amp/: Complete Amp routing module - sdk/api/httpx/: HTTP utilities for gzip/transport - 94.6% test coverage with 34 comprehensive test cases - Clean integration minimizes merge conflict risk Security: - Management routes restricted to localhost by default - Configurable via amp-restrict-management-to-localhost - Prevents drive-by browser attacks on user data This provides a production-ready foundation for Amp CLI integration while maintaining clean separation from upstream code for easy rebasing. Amp-Thread-ID: https://ampcode.com/threads/T-9e2befc5-f969-41c6-890c-5b779d58cf18	2025-11-19 18:23:17 -07:00
Luis Pater	618511ff67	Merge pull request #280 from ben-vargas/feat-enable-gemini-3-cli feat: enable Gemini 3 Pro Preview with OAuth support	2025-11-20 08:46:57 +08:00
Ben Vargas	0ff094b87f	fix(executor): prevent streaming on failed response when no fallback Fix critical bug where ExecuteStream would create a streaming channel from a failed (non-2xx) response after exhausting all retries with no fallback models available. When retries were exhausted on the last model, the code would break from the inner loop but fall through to streaming channel creation (line 401), immediately returning at line 461. This made the error handling code at lines 464-471 unreachable, causing clients to receive an empty/closed stream instead of a proper error response. Solution: Check if httpResp is non-2xx before creating the streaming channel. If failed, continue the outer loop to reach error handling. Identified by: codex-bot review Ref: https://github.com/router-for-me/CLIProxyAPI/pull/280#pullrequestreview-3484560423	2025-11-19 13:14:40 -07:00
Ben Vargas	ed23472d94	fix(executor): prevent streaming from 429 response when fallback available Fix critical bug where ExecuteStream would create a streaming channel using a 429 error response instead of continuing to the next fallback model after exhausting retries. When 429 retries were exhausted and a fallback model was available, the inner retry loop would break but immediately fall through to the streaming channel creation, attempting to stream from the failed 429 response instead of trying the next model. Solution: Add shouldContinueToNextModel flag to explicitly skip the streaming logic and continue the outer model loop when appropriate. Identified by: codex-bot review Ref: https://github.com/router-for-me/CLIProxyAPI/pull/280#pullrequestreview-3484479106	2025-11-19 13:05:38 -07:00
Ben Vargas	ede4471b84	feat(translator): add default thinkingConfig for gemini-3-pro-preview Match official Gemini CLI behavior by always sending default thinkingConfig when client doesn't specify reasoning parameters. - Set thinkingBudget=-1 (dynamic) for gemini-3-pro-preview - Set include_thoughts=true to return thinking process - Apply to both /v1/chat/completions and /v1/responses endpoints - See: ai-gemini-cli/packages/core/src/config/defaultModelConfigs.ts	2025-11-19 12:47:39 -07:00
Ben Vargas	6a3de3a89c	feat(executor): add intelligent retry logic for 429 rate limits Implement Google RetryInfo.retryDelay support for handling 429 rate limit errors. Retries same model up to 3 times using exact delays from Google's API before trying fallback models. - Add parseRetryDelay() to extract Google's retry guidance - Implement inner retry loop in Execute() and ExecuteStream() - Context-aware waiting with cancellation support - Cap delays at 60s maximum for safety	2025-11-19 12:47:39 -07:00
Ben Vargas	782bba0bc4	feat(registry): enable gemini-3-pro-preview for gemini-cli provider Add gemini-3-pro-preview model to GetGeminiCLIModels() to make it available for OAuth-based Gemini CLI users, matching the model already available in AI Studio provider. Model spec: - ID: gemini-3-pro-preview - Version: 3.0 - Input: 1M tokens - Output: 64K tokens - Thinking: 128-32K tokens (dynamic)	2025-11-19 12:47:39 -07:00
Luis Pater	bf116b68f8	feat(registry): add GPT-5.1 Codex Max model definitions and support - Introduced `gpt-5.1-codex-max` variants to model definitions (`low`, `medium`, `high`, `xhigh`). - Updated executor logic to map effort levels for Codex Max models. - Added `lastCodexMaxPrompt` processing for `gpt-5.1-codex-max` prompts. - Defined instructions for `gpt-5.1-codex-max` in a new file: `codex_instructions/gpt-5.1-codex-max_prompt.md`. v6.3.57	2025-11-20 03:12:22 +08:00
Luis Pater	cc3cf09c00	feat(auth): add AuthIndex for diagnostics and ensure usage recording v6.3.56	2025-11-19 22:02:40 +08:00
Luis Pater	9acfbcc2a0	Merge pull request #275 from router-for-me/iflow Iflow v6.3.55	2025-11-19 20:44:54 +08:00
hkfires	b285b07986	fix(iflow): adjust auth filename email sanitization	2025-11-19 19:50:06 +08:00
Luis Pater	c40e00526b	Merge pull request #274 from router-for-me/log fix: detect HTML error bodies without text/html content type	2025-11-19 17:40:06 +08:00
hkfires	8a33f3ef69	fix: detect HTML error bodies without text/html content type	2025-11-19 14:45:33 +08:00
Luis Pater	7a8e00fcea	fix(translator): handle missing parameters in Gemini tool schema gracefully v6.3.54	2025-11-19 13:19:46 +08:00
Luis Pater	89771216a1	feat(translator): add ThoughtSignature handling in Gemini request transformations v6.3.53	2025-11-19 11:34:13 +08:00
Luis Pater	14ddfd4b79	Merge pull request #270 from router-for-me/iflow feat(auth): add iFlow cookie-based authentication support	2025-11-19 01:54:34 +08:00
Luis Pater	567227f35f	Merge pull request #268 from router-for-me/tools fix: use underscore suffix in short name mapping	2025-11-19 01:43:41 +08:00
Luis Pater	17016ae6a5	feat(registry): add Gemini 3 Pro Preview model definition v6.3.52	2025-11-18 23:48:21 +08:00
Luis Pater	01b7b60901	feat(registry): add Gemini 3 Pro Preview model definition	2025-11-18 23:46:58 +08:00
hkfires	b52a5cc066	feat(auth): add iFlow cookie-based authentication support	2025-11-18 22:35:35 +08:00
hkfires	1ba057112a	fix: use underscore suffix in short name mapping Replace the "~<n>" suffix with "_<n>" when generating unique short names in codex translators (Claude, Gemini, OpenAI chat). This avoids using a special character in identifiers, improving compatibility with downstream APIs while preserving length constraints.	2025-11-18 16:59:25 +08:00
Luis Pater	23a7633e6d	fix(registry): update Thinking parameters and replace Gemini-3 Preview with Gemini-2.5 Flash Lite v6.3.51	2025-11-18 11:51:52 +08:00
Luis Pater	e5e985978d	Fixed: #263 fix(translator): remove input_examples from tool schema in Gemini-Claude requests v6.3.50	2025-11-18 11:27:48 +08:00
Luis Pater	db2d22c978	fix(runtime): simplify scanner buffer allocation in executor implementations	2025-11-18 10:59:49 +08:00
Luis Pater	1c815c58a6	fix(translator): simplify string handling in Gemini responses v6.3.49	2025-11-16 19:02:27 +08:00
Luis Pater	4eab141410	feat(translator): add support for reasoning/thinking content blocks in OpenAI-Claude and Gemini responses v6.3.48	2025-11-16 17:37:39 +08:00
Luis Pater	5937b8e429	Fixed: #260 fix(translator): handle simple string input conversion in Gemini responses v6.3.47	2025-11-16 13:30:11 +08:00
Luis Pater	9875565339	fix(claude translator): ensure default token counts when usage data is missing v6.3.46	2025-11-16 13:18:21 +08:00
Luis Pater	faa483b57d	Merge pull request #257 from lollipopkit/main fix(claude translator): guard tool schema properties v6.3.45	2025-11-16 12:19:38 +08:00
Luis Pater	f0711be302	fix(auth): prevent access to removed credentials lingering in memory Add logic to avoid exposing credentials that have been removed from disk but still persist in memory. Ensure `runtimeOnly` checks and proper handling of disabled or removed authentication states. v6.3.44	2025-11-16 12:12:24 +08:00
Luis Pater	1d0f0301b4	refactor(api/config): centralize legacy OpenAI compatibility key migration Introduce `migrateLegacyOpenAICompatibilityKeys` to streamline and reuse the normalization of OpenAI compatibility entries. Remove redundant loops and enhance maintainability for compatibility key handling. Add cleanup for legacy `api-keys` in YAML configuration during persistence.	2025-11-16 11:39:35 +08:00
lollipopkit🏳️‍⚧️	c73b3fa43b	fix(claude translator): guard tool schema properties	2025-11-15 19:14:13 +08:00
Luis Pater	772fa69515	Fixed: #254 feat(registry): add Kimi-K2-Thinking model to model definitions v6.3.43	2025-11-14 21:20:54 +08:00
Luis Pater	1ccb01631d	refactor(runtime): centralize reasoning effort logic for GPT models Extract reasoning effort mapping into a reusable function `setReasoningEffortByAlias` to reduce redundancy and improve maintainability. Introduce support for the "gpt-5.1-none" variant in the registry and runtime executor. v6.3.42	2025-11-14 17:24:40 +08:00
Luis Pater	1ede1347fa	Merge pull request #249 from ben-vargas/fix-gpt5-1-reasoning fix(runtime): remove gpt-5.1 minimal effort variant	2025-11-14 17:04:27 +08:00
Ben Vargas	cfbaed0e90	fix(runtime): remove gpt-5.1 minimal effort variant Stop advertising and mapping the unsupported gpt-5.1-minimal variant in the model registry and Codex executor, and align bare gpt-5.1 requests to use medium reasoning effort like Codex CLI while preserving minimal for gpt-5.	2025-11-13 19:43:52 -07:00
Luis Pater	cf9b9be7ea	feat(runtime): extend executor support for GPT-5.1 Codex and variants Expand executor logic to handle GPT-5.1 Codex family and its variants, including reasoning effort configurations for minimal, low, medium, and high levels. Ensure proper mapping of models to payload parameters. v6.3.41	2025-11-14 08:08:25 +08:00
Luis Pater	aa57f3237a	feat(instructions): add detailed agent behavior guidelines for Codex CLI Introduce comprehensive agent instruction documentation (`gpt_5_1_prompt.md`) for Codex CLI. Specify agent behavior, personality, planning requirements, task execution, sandboxing rules, and validation processes to standardize interactions and improve usability. v6.3.40	2025-11-14 06:51:54 +08:00
Luis Pater	fcd98f4f9b	feat(runtime): add payload configuration support for executors Introduce `PayloadConfig` in the configuration to define default and override rules for modifying payload parameters. Implement `applyPayloadConfig` and `applyPayloadConfigWithRoot` to apply these rules across various executors, ensuring consistent parameter handling for different models and protocols. Update all relevant executors to utilize this functionality. v6.3.39	2025-11-13 23:27:40 +08:00
Luis Pater	75b57bc112	Fixed: #246 feat(runtime): add support for GPT-5.1 models and variants Introduce GPT-5.1 model family, including minimal, low, medium, high, Codex, and Codex Mini variants. Update tokenization and reasoning effort handling to accommodate new models in executor and registry. v6.3.38	2025-11-13 17:42:19 +08:00

1 2 3 4 5 ...

734 Commits