Commit Graph

62 Commits

  • Fix image limits test to use realistic payload sizes
    Previous test used compressed 8k images (0.01MB) which was meaningless.
    Now tests with actual large noise images that don't compress.
    
    Realistic payload limits discovered:
    - Anthropic: 6 x 3MB = ~18MB total (not 32MB as documented)
    - OpenAI: 2 x 15MB = ~30MB total
    - Gemini: 10 x 20MB = ~200MB total (very permissive)
    - Mistral: 4 x 10MB = ~40MB total
    - xAI: 1 x 20MB (strict request size limit)
    - Groq: 5 x 5760px images (5 image + pixel limit)
    - zAI: 2 x 15MB = ~30MB (50MB request limit)
    - OpenRouter: 2 x 5MB = ~10MB total
    
    Also fixed GEMINI_API_KEY env var (was GOOGLE_API_KEY).
    
    Related to #120
  • Update image limits test with comprehensive 8k stress test results
    Tested max 8kx8k images per provider:
    - Anthropic: 100 (explicit limit, fails at 101)
    - OpenAI: 100-200 (100 works, 200 times out)
    - Mistral: 8 (explicit limit, fails at 9)
    - xAI: 100-150 (100 works, 150 times out)
    - Groq: 0 (8k exceeds 33M pixel limit)
    - zAI: 400 (context window limited at 500)
    - OpenRouter: 40 (context window limited at 50)
    - Gemini: untested (no API key in test env)
    
    Key finding: Anthropic's 'many images' rule does NOT cause API errors.
    100 x 8kx8k images work fine. Anthropic likely auto-resizes internally.
    
    Related to #120
  • Add comprehensive image limits test suite for all vision-capable providers
    Tests max image count, size, dimensions, and 8k stress test for:
    - Anthropic, OpenAI, Gemini, Mistral, OpenRouter, xAI, Groq, zAI
    
    Key finding: Anthropic's 'many images' rule (>20 images = 2000px max)
    does NOT cause API errors. 100 x 8k images work fine. Anthropic likely
    auto-resizes internally.
    
    Related to #120
  • ai: add image limits test suite
    Tests provider-specific image limitations across all supported providers:
    - Maximum number of images in context
    - Maximum image size (bytes)
    - Maximum image dimensions
    
    Discovered limits (Dec 2025):
    - Anthropic: 100 images, 5MB per image, 8000px max dimension
    - OpenAI: 500 images, >=25MB per image
    - Gemini: ~2500 images, >=40MB per image
    - Mistral: 8 images, ~15MB per image
    - OpenRouter: ~40 images (context limited), ~15MB per image
  • Fix branch selector for single message and --no-session mode
    - Allow branch selector to open with single user message (changed <= 1 to === 0 check)
    - Support in-memory branching for --no-session mode (no files created)
    - Add isEnabled() getter to SessionManager
    - Update sessionFile getter to return null when sessions disabled
    - Update SessionSwitchEvent types to allow null session files
    - Add branching tests for single message and --no-session scenarios
    
    fixes #163
  • Fix Mistral 400 errors after aborted assistant messages
    - Skip empty assistant messages (no content, no tool calls) to avoid
      Mistral's 'Assistant message must have either content or tool_calls'
      error
    - Remove synthetic assistant bridge message after tool results (Mistral
      no longer requires this as of Dec 2024)
    - Add test for empty assistant message handling
    
    Follow-up to #165
  • Add Mistral as AI provider
    - Add Mistral to KnownProvider type and model generation
    - Implement Mistral-specific compat handling in openai-completions:
      - requiresToolResultName: tool results need name field
      - requiresAssistantAfterToolResult: synthetic assistant message between tool/user
      - requiresThinkingAsText: thinking blocks as <thinking> text
      - requiresMistralToolIds: tool IDs must be exactly 9 alphanumeric chars
    - Add MISTRAL_API_KEY environment variable support
    - Add Mistral tests across all test files
    - Update documentation (README, CHANGELOG) for both ai and coding-agent packages
    - Remove client IDs from gemini.md, reference upstream source instead
    
    Closes #165
  • Simplify compaction: remove proactive abort, use Agent.continue() for retry
    - Add agentLoopContinue() to pi-ai for resuming from existing context
    - Add Agent.continue() method and transport.continue() interface
    - Simplify AgentSession compaction to two cases: overflow (auto-retry) and threshold (no retry)
    - Remove proactive mid-turn compaction abort
    - Merge turn prefix summary into main summary
    - Add isCompacting property to AgentSession and RPC state
    - Block input during compaction in interactive mode
    - Show compaction count on session resume
    - Rename RPC.md to rpc.md for consistency
    
    Related to #128
  • Add xhigh thinking level for OpenAI codex-max models
    - Add 'xhigh' to ThinkingLevel type in ai and agent packages
    - Map xhigh to reasoning_effort: 'max' for OpenAI providers
    - Add thinkingXhigh color token to theme schema and built-in themes
    - Show xhigh option only when using codex-max models
    - Update CHANGELOG for both ai and coding-agent packages
    
    closes #143
  • Implement tool result truncation with actionable notices (#134)
    - read: actionable notices with offset for continuation
      - First line > 30KB: return empty + bash command suggestion
      - Hit limit: '[Showing lines X-Y of Z. Use offset=N to continue]'
    
    - bash: tail truncation with temp file
      - Notice includes line range + temp file path
      - Edge case: last line > 30KB shows partial
    
    - grep: pre-truncate match lines to 500 chars
      - '[... truncated]' suffix on long lines
      - Notice for match limit and line truncation
    
    - find/ls: result/entry limit notices
      - '[N results limit reached. Use limit=M for more]'
    
    - All notices now in text content (LLM sees them)
    - TUI simplified (notices render as part of output)
    - Never return partial lines (except bash edge case)
  • Add totalTokens field to Usage type
    - Added totalTokens field to Usage interface in pi-ai
    - Anthropic: computed as input + output + cacheRead + cacheWrite
    - OpenAI/Google: uses native total_tokens/totalTokenCount
    - Fixed openai-completions to compute totalTokens when reasoning tokens present
    - Updated calculateContextTokens() to use totalTokens field
    - Added comprehensive test covering 13 providers
    
    fixes #130
  • Add context overflow detection utilities
    Extract overflow detection logic into reusable utilities:
    - isContextOverflowError() to detect overflow from error messages
    - isContextOverflowFromUsage() to detect overflow from token usage
    - Patterns for Anthropic, OpenAI, Google, xAI, Groq, OpenRouter, llama.cpp, LM Studio
    
    Fixes #129
  • Add image support in tool results across all providers
    Tool results now use content blocks and can include both text and images.
    All providers (Anthropic, Google, OpenAI Completions, OpenAI Responses)
    correctly pass images from tool results to LLMs.
    
    - Update ToolResultMessage type to use content blocks
    - Add placeholder text for image-only tool results in Google/Anthropic
    - OpenAI providers send tool result + follow-up user message with images
    - Fix Anthropic JSON parsing for empty tool arguments
    - Add comprehensive tests for image-only and text+image tool results
    - Update README with tool result content blocks API
  • Fix token statistics on abort for Anthropic provider
    - Add handling for message_start event to capture initial token usage
    - Fix message_delta to use assignment (=) instead of addition (+=)
      since Anthropic sends cumulative token counts, not incremental
    - Add comprehensive tests for all providers (Google, OpenAI Completions,
      OpenAI Responses, Anthropic)
    - Document OpenAI limitation: token stats only available at stream end
    
    Fixes issue where aborted streams had zero token counts despite
    Anthropic sending input tokens in the initial message_start event.
  • Add Unicode surrogate sanitization for all providers
    Fixes issue where unpaired Unicode surrogates in tool results cause JSON serialization errors in API providers, particularly Anthropic.
    
    - Add sanitizeSurrogates() utility function to remove unpaired surrogates
    - Apply sanitization in all provider convertMessages() functions:
      - User message text content (string and text blocks)
      - Assistant message text and thinking blocks
      - Tool result output
      - System prompts
    - Valid emoji (properly paired surrogates) are preserved
    - Add comprehensive test suite covering all 8 providers
    
    Previously only Google and Groq handled unpaired surrogates correctly.
    Now all providers (Anthropic, OpenAI Completions/Responses, Google, xAI, Groq, Cerebras, zAI) sanitize text before API submission.
  • Add ToolRenderResult interface for custom tool rendering
    - Changed ToolRenderer return type from TemplateResult to ToolRenderResult
    - ToolRenderResult = { content: TemplateResult, isCustom: boolean }
    - isCustom: true = no card wrapper, false = wrap in card
    - Updated all existing tool renderers to return new format
    - Updated Messages.ts to handle custom rendering
    
    This enables tools to render without default card chrome when needed.
  • refactor(ai): improve error handling and stop reason types
    - Add 'aborted' as a distinct stop reason separate from 'error'
    - Change AssistantMessage.error to errorMessage for clarity
    - Update error event to include reason field ('error' | 'aborted')
    - Map provider-specific safety/refusal reasons to 'error' stop reason
    - Reorganize utility functions into utils/ directory
    - Rename agent.ts to agent-loop.ts for better clarity
    - Fix error handling in all providers to properly distinguish abort from error
  • feat(ai): add partial JSON parsing for streaming tool calls
    - Added partial-json package for parsing incomplete JSON during streaming
    - Tool call arguments now contain partially parsed JSON during toolcall_delta events
    - Enables progressive UI updates (e.g., showing file paths before content is complete)
    - Arguments are always valid objects (minimum empty {}), never undefined
    - Full validation still occurs at toolcall_end when arguments are complete
    - Updated all providers (Anthropic, OpenAI Completions/Responses) to use parseStreamingJson
    - Added comprehensive documentation and examples in README
    - Added test to verify arguments are always defined during streaming
  • Replace Zod with TypeBox for schema validation
    - Switch from Zod to TypeBox for tool parameter schemas
    - TypeBox schemas can be serialized/deserialized as JSON
    - Use AJV for runtime validation instead of Zod's parse
    - Add StringEnum helper for Google API compatibility (avoids anyOf/const patterns)
    - Export Type and Static from main package for convenience
    - Update all tests and documentation to reflect TypeBox usage
  • feat(ai): Implement Zod-based tool validation and improve Agent API
    - Replace JSON Schema with Zod schemas for tool parameter definitions
    - Add runtime validation for all tool calls at provider level
    - Create shared validation module with detailed error formatting
    - Update Agent API with comprehensive event system
    - Add agent tests with calculator tool for multi-turn execution
    - Add abort test to verify proper handling of aborted requests
    - Update documentation with detailed event flow examples
    - Rename generate.ts to stream.ts for clarity
  • feat(ai): Add zAI provider support
    - Add 'zai' as a KnownProvider type
    - Add ZAI_API_KEY environment variable mapping
    - Generate 4 zAI models (glm-4.5-air, glm-4.5v, etc.) using anthropic-messages API
    - Add comprehensive test coverage for zAI provider in generate.test.ts and empty.test.ts
    - Models support reasoning/thinking capabilities and tool calling
  • fix(ai): Sanitize tool call IDs for Anthropic API compatibility
    - Anthropic API requires tool call IDs to match pattern ^[a-zA-Z0-9_-]+$
    - OpenAI Responses API generates IDs with pipe character (|) which breaks Anthropic
    - Added sanitizeToolCallId() to replace invalid characters with underscores
    - Fixes cross-provider handoffs from OpenAI Responses to Anthropic
    - Added test to verify the fix works
  • Massive refactor of API
    - Switch to function based API
    - Anthropic SDK style async generator
    - Fully typed with escape hatches for custom models
  • feat(ai): Add new streaming generate API with AsyncIterable interface
    - Implement QueuedGenerateStream class that extends AsyncIterable with finalMessage() method
    - Add new types: GenerateStream, GenerateOptions, GenerateOptionsUnified, GenerateFunction
    - Create generateAnthropic function-based implementation replacing class-based approach
    - Add comprehensive test suite for the new generate API
    - Support streaming events with text, thinking, and tool call deltas
    - Map ReasoningEffort to provider-specific options
    - Include apiKey in options instead of constructor parameter
  • test(ai): Add empty assistant message tests
    - Test providers handling empty assistant messages in conversation flow
    - Pattern: user message -> empty assistant -> user message
    - All providers handle empty assistant messages gracefully
    - Tests ensure providers can continue conversation after empty response
  • test(ai): Add empty message tests for all providers
    - Test handling of empty content arrays
    - Test handling of empty string content
    - Test handling of whitespace-only content
    - All providers handle these edge cases gracefully
  • feat(ai): Fetch Anthropic, Google, and OpenAI models from models.dev instead of OpenRouter
    - Updated generate-models.ts to fetch these providers directly from models.dev API
    - OpenRouter now only used for xAI and other third-party providers
    - Fixed test model IDs to match new model names from models.dev
    - Removed unused import from google.ts
  • feat(ai): Add cross-provider message handoff support
    - Add transformMessages utility to handle cross-provider compatibility
    - Convert thinking blocks to <thinking> tagged text when switching providers
    - Preserve native thinking blocks when staying with same provider/model
    - Add comprehensive handoff tests verifying all provider combinations
    - Fix OpenAI Completions to return partial results on abort
    - Update tool call ID format for Anthropic compatibility
    - Document cross-provider handoff capabilities in README
  • refactor(ai): Update API to support partial results on abort
    - Anthropic, Google, and OpenAI Responses providers now return partial results when aborted
    - Restructured streaming to accumulate content blocks incrementally
    - Prevents submission of thinking/toolCall blocks from aborted completions in multi-turn conversations
    - Makes UI development easier by providing partial content even when requests are interrupted
  • feat(ai): Add start event emission to all providers
    - Emit start event with model and provider info after creating stream
    - Add abort signal tests for all providers
    - Update README abort signal section to reflect non-throwing API
    - Fix model references in README examples
  • fix(ai): Fix OpenAI Responses provider multi-turn conversation support
    - Collect complete output items during streaming instead of building blocks incrementally
    - Handle reasoning summary parts with proper newline separation
    - Support refusal content in message outputs
    - Preserve full reasoning items and message IDs for multi-turn resubmission
    - Emit proper streaming events for text and thinking deltas
  • refactor(ai): Update API to support multiple thinking and text blocks
    BREAKING CHANGE: AssistantMessage now uses content array instead of separate fields
    - Changed AssistantMessage.content from string to array of content blocks
    - Removed separate thinking, toolCalls, and signature fields
    - Content blocks can be TextContent, ThinkingContent, or ToolCall types
    - Updated streaming events to include start/end events for text and thinking
    - Fixed multiTurn test to handle new content structure
    
    Note: Currently only Anthropic provider is updated to work with new API
    Other providers need to be updated to match the new interface
  • test(ai): Add image input test for Anthropic Haiku 3.5
    - Added image handling test for Claude 3.5 Haiku
    - Ensures vision capabilities are properly tested
  • fix(ai): Fix OpenAI Responses provider multi-turn conversation support
    - Added contentSignature tracking for assistant messages
    - Fixed message format in convertToResponsesFormat (output_text instead of input_text)
    - Properly preserve message IDs for multi-turn conversations
    - Added proper ResponseOutputMessage type satisfaction
    - Updated tests to cover more providers and multi-turn scenarios
  • feat(ai): Add image input tests for vision-capable models
    - Added image tests to OpenAI Completions (gpt-4o-mini)
    - Added image tests to Anthropic (claude-sonnet-4-0)
    - Added image tests to Google (gemini-2.5-flash)
    - Tests verify models can process and describe the red circle test image
  • refactor(ai): Update LLM implementations to use Model objects
    - LLM constructors now take Model objects instead of string IDs
    - Added provider field to AssistantMessage interface
    - Updated getModel function with type-safe model ID autocomplete
    - Fixed Anthropic model ID mapping for proper API aliases
    - Added baseUrl to Model interface for provider-specific endpoints
    - Updated all tests to use getModel for model instantiation
    - Removed deprecated models.json in favor of generated models