Commit Graph

46 Commits

  • Add image support in tool results across all providers
    Tool results now use content blocks and can include both text and images.
    All providers (Anthropic, Google, OpenAI Completions, OpenAI Responses)
    correctly pass images from tool results to LLMs.
    
    - Update ToolResultMessage type to use content blocks
    - Add placeholder text for image-only tool results in Google/Anthropic
    - OpenAI providers send tool result + follow-up user message with images
    - Fix Anthropic JSON parsing for empty tool arguments
    - Add comprehensive tests for image-only and text+image tool results
    - Update README with tool result content blocks API
  • Fix token statistics on abort for Anthropic provider
    - Add handling for message_start event to capture initial token usage
    - Fix message_delta to use assignment (=) instead of addition (+=)
      since Anthropic sends cumulative token counts, not incremental
    - Add comprehensive tests for all providers (Google, OpenAI Completions,
      OpenAI Responses, Anthropic)
    - Document OpenAI limitation: token stats only available at stream end
    
    Fixes issue where aborted streams had zero token counts despite
    Anthropic sending input tokens in the initial message_start event.
  • Add Unicode surrogate sanitization for all providers
    Fixes issue where unpaired Unicode surrogates in tool results cause JSON serialization errors in API providers, particularly Anthropic.
    
    - Add sanitizeSurrogates() utility function to remove unpaired surrogates
    - Apply sanitization in all provider convertMessages() functions:
      - User message text content (string and text blocks)
      - Assistant message text and thinking blocks
      - Tool result output
      - System prompts
    - Valid emoji (properly paired surrogates) are preserved
    - Add comprehensive test suite covering all 8 providers
    
    Previously only Google and Groq handled unpaired surrogates correctly.
    Now all providers (Anthropic, OpenAI Completions/Responses, Google, xAI, Groq, Cerebras, zAI) sanitize text before API submission.
  • Add ToolRenderResult interface for custom tool rendering
    - Changed ToolRenderer return type from TemplateResult to ToolRenderResult
    - ToolRenderResult = { content: TemplateResult, isCustom: boolean }
    - isCustom: true = no card wrapper, false = wrap in card
    - Updated all existing tool renderers to return new format
    - Updated Messages.ts to handle custom rendering
    
    This enables tools to render without default card chrome when needed.
  • refactor(ai): improve error handling and stop reason types
    - Add 'aborted' as a distinct stop reason separate from 'error'
    - Change AssistantMessage.error to errorMessage for clarity
    - Update error event to include reason field ('error' | 'aborted')
    - Map provider-specific safety/refusal reasons to 'error' stop reason
    - Reorganize utility functions into utils/ directory
    - Rename agent.ts to agent-loop.ts for better clarity
    - Fix error handling in all providers to properly distinguish abort from error
  • feat(ai): add partial JSON parsing for streaming tool calls
    - Added partial-json package for parsing incomplete JSON during streaming
    - Tool call arguments now contain partially parsed JSON during toolcall_delta events
    - Enables progressive UI updates (e.g., showing file paths before content is complete)
    - Arguments are always valid objects (minimum empty {}), never undefined
    - Full validation still occurs at toolcall_end when arguments are complete
    - Updated all providers (Anthropic, OpenAI Completions/Responses) to use parseStreamingJson
    - Added comprehensive documentation and examples in README
    - Added test to verify arguments are always defined during streaming
  • Replace Zod with TypeBox for schema validation
    - Switch from Zod to TypeBox for tool parameter schemas
    - TypeBox schemas can be serialized/deserialized as JSON
    - Use AJV for runtime validation instead of Zod's parse
    - Add StringEnum helper for Google API compatibility (avoids anyOf/const patterns)
    - Export Type and Static from main package for convenience
    - Update all tests and documentation to reflect TypeBox usage
  • feat(ai): Implement Zod-based tool validation and improve Agent API
    - Replace JSON Schema with Zod schemas for tool parameter definitions
    - Add runtime validation for all tool calls at provider level
    - Create shared validation module with detailed error formatting
    - Update Agent API with comprehensive event system
    - Add agent tests with calculator tool for multi-turn execution
    - Add abort test to verify proper handling of aborted requests
    - Update documentation with detailed event flow examples
    - Rename generate.ts to stream.ts for clarity
  • feat(ai): Add zAI provider support
    - Add 'zai' as a KnownProvider type
    - Add ZAI_API_KEY environment variable mapping
    - Generate 4 zAI models (glm-4.5-air, glm-4.5v, etc.) using anthropic-messages API
    - Add comprehensive test coverage for zAI provider in generate.test.ts and empty.test.ts
    - Models support reasoning/thinking capabilities and tool calling
  • fix(ai): Sanitize tool call IDs for Anthropic API compatibility
    - Anthropic API requires tool call IDs to match pattern ^[a-zA-Z0-9_-]+$
    - OpenAI Responses API generates IDs with pipe character (|) which breaks Anthropic
    - Added sanitizeToolCallId() to replace invalid characters with underscores
    - Fixes cross-provider handoffs from OpenAI Responses to Anthropic
    - Added test to verify the fix works
  • Massive refactor of API
    - Switch to function based API
    - Anthropic SDK style async generator
    - Fully typed with escape hatches for custom models
  • feat(ai): Add new streaming generate API with AsyncIterable interface
    - Implement QueuedGenerateStream class that extends AsyncIterable with finalMessage() method
    - Add new types: GenerateStream, GenerateOptions, GenerateOptionsUnified, GenerateFunction
    - Create generateAnthropic function-based implementation replacing class-based approach
    - Add comprehensive test suite for the new generate API
    - Support streaming events with text, thinking, and tool call deltas
    - Map ReasoningEffort to provider-specific options
    - Include apiKey in options instead of constructor parameter
  • test(ai): Add empty assistant message tests
    - Test providers handling empty assistant messages in conversation flow
    - Pattern: user message -> empty assistant -> user message
    - All providers handle empty assistant messages gracefully
    - Tests ensure providers can continue conversation after empty response
  • test(ai): Add empty message tests for all providers
    - Test handling of empty content arrays
    - Test handling of empty string content
    - Test handling of whitespace-only content
    - All providers handle these edge cases gracefully
  • feat(ai): Fetch Anthropic, Google, and OpenAI models from models.dev instead of OpenRouter
    - Updated generate-models.ts to fetch these providers directly from models.dev API
    - OpenRouter now only used for xAI and other third-party providers
    - Fixed test model IDs to match new model names from models.dev
    - Removed unused import from google.ts
  • feat(ai): Add cross-provider message handoff support
    - Add transformMessages utility to handle cross-provider compatibility
    - Convert thinking blocks to <thinking> tagged text when switching providers
    - Preserve native thinking blocks when staying with same provider/model
    - Add comprehensive handoff tests verifying all provider combinations
    - Fix OpenAI Completions to return partial results on abort
    - Update tool call ID format for Anthropic compatibility
    - Document cross-provider handoff capabilities in README
  • refactor(ai): Update API to support partial results on abort
    - Anthropic, Google, and OpenAI Responses providers now return partial results when aborted
    - Restructured streaming to accumulate content blocks incrementally
    - Prevents submission of thinking/toolCall blocks from aborted completions in multi-turn conversations
    - Makes UI development easier by providing partial content even when requests are interrupted
  • feat(ai): Add start event emission to all providers
    - Emit start event with model and provider info after creating stream
    - Add abort signal tests for all providers
    - Update README abort signal section to reflect non-throwing API
    - Fix model references in README examples
  • fix(ai): Fix OpenAI Responses provider multi-turn conversation support
    - Collect complete output items during streaming instead of building blocks incrementally
    - Handle reasoning summary parts with proper newline separation
    - Support refusal content in message outputs
    - Preserve full reasoning items and message IDs for multi-turn resubmission
    - Emit proper streaming events for text and thinking deltas
  • refactor(ai): Update API to support multiple thinking and text blocks
    BREAKING CHANGE: AssistantMessage now uses content array instead of separate fields
    - Changed AssistantMessage.content from string to array of content blocks
    - Removed separate thinking, toolCalls, and signature fields
    - Content blocks can be TextContent, ThinkingContent, or ToolCall types
    - Updated streaming events to include start/end events for text and thinking
    - Fixed multiTurn test to handle new content structure
    
    Note: Currently only Anthropic provider is updated to work with new API
    Other providers need to be updated to match the new interface
  • test(ai): Add image input test for Anthropic Haiku 3.5
    - Added image handling test for Claude 3.5 Haiku
    - Ensures vision capabilities are properly tested
  • fix(ai): Fix OpenAI Responses provider multi-turn conversation support
    - Added contentSignature tracking for assistant messages
    - Fixed message format in convertToResponsesFormat (output_text instead of input_text)
    - Properly preserve message IDs for multi-turn conversations
    - Added proper ResponseOutputMessage type satisfaction
    - Updated tests to cover more providers and multi-turn scenarios
  • feat(ai): Add image input tests for vision-capable models
    - Added image tests to OpenAI Completions (gpt-4o-mini)
    - Added image tests to Anthropic (claude-sonnet-4-0)
    - Added image tests to Google (gemini-2.5-flash)
    - Tests verify models can process and describe the red circle test image
  • refactor(ai): Update LLM implementations to use Model objects
    - LLM constructors now take Model objects instead of string IDs
    - Added provider field to AssistantMessage interface
    - Updated getModel function with type-safe model ID autocomplete
    - Fixed Anthropic model ID mapping for proper API aliases
    - Added baseUrl to Model interface for provider-specific endpoints
    - Updated all tests to use getModel for model instantiation
    - Removed deprecated models.json in favor of generated models
  • fix(ai): Deduplicate models and add Anthropic aliases
    - Add proper Anthropic model aliases (claude-opus-4-1, claude-sonnet-4-0, etc.)
    - Deduplicate models when same ID appears in both models.dev and OpenRouter
    - models.dev takes priority over OpenRouter for duplicate IDs
    - Fix test to use correct claude-3-5-haiku-latest alias
    - Reduces Anthropic models from 11 to 10 (removed duplicate)
  • refactor(ai): Implement unified model system with type-safe createLLM
    - Add Model interface to types.ts with normalized structure
    - Create type-safe generic createLLM function with provider-specific model constraints
    - Generate models from OpenRouter API and models.dev data
    - Strip provider prefixes for direct providers (google, openai, anthropic, xai)
    - Keep full model IDs for OpenRouter-proxied models
    - Clean separation: types.ts (Model interface), models.ts (factory logic), models.generated.ts (data)
    - Remove old model scripts and unused dependencies
    - Rename GeminiLLM to GoogleLLM for consistency
    - Add tests for new providers (xAI, Groq, Cerebras, OpenRouter)
    - Support 181 tool-capable models across 7 providers with full type safety
  • feat(ai): Migrate tests to Vitest and add provider test coverage
    - Switch from Node.js test runner to Vitest for better DX
    - Add test suites for Grok, Groq, Cerebras, and OpenRouter providers
    - Add Ollama test suite with automatic server lifecycle management
    - Include thinking mode and multi-turn tests for all providers
    - Remove example files (consolidated into test suite)
    - Add VS Code test configuration
  • fix(ai): Improve ModelInfo types based on actual data structure
    - Remove catch-all [key: string]: any from ModelInfo
    - Make all required fields non-optional (attachment, reasoning, etc.)
    - Add proper union types for modalities (text, image, audio, video, pdf)
    - Mark only cost and knowledge fields as optional
    - Export ModalityInput and ModalityOutput types
  • feat(ai): Add models.dev data integration
    - Add models script to download latest model information
    - Create models.ts module to query model capabilities
    - Include models.json in package distribution
    - Export utilities to check model features (reasoning, tools)
    - Update build process to copy models.json to dist
  • feat(ai): Add OpenAI-compatible provider examples for multiple services
    - Add examples for Cerebras, Groq, Ollama, and OpenRouter
    - Update OpenAI Completions provider to handle base URL properly
    - Simplify README formatting
    - All examples use the same OpenAICompletionsLLM provider with different base URLs
  • test(ai): Add comprehensive E2E tests for all AI providers
    - Add multi-turn test to verify thinking and tool calling work together
    - Test thinkingSignature handling for proper multi-turn context
    - Fix Gemini provider to generate base64 thinkingSignature when needed
    - Handle multiple rounds of tool calls in tests (Gemini behavior)
    - Make thinking tests more robust for model-dependent behavior
    - All 18 tests passing across 4 providers
  • feat(ai): Add proper thinking support for Gemini 2.5 models
    - Added thinkingConfig with includeThoughts and thinkingBudget support
    - Use part.thought boolean flag to detect thinking content per API docs
    - Capture and preserve thought signatures for multi-turn function calling
    - Added supportsThinking() check for Gemini 2.5 series models
    - Updated example to demonstrate thinking configuration
    - Handle SDK type limitations with proper type assertions
  • feat(ai): Implement Gemini provider with streaming and tool support
    - Added GeminiLLM provider implementation with GoogleGenerativeAI SDK
    - Supports streaming with text/thinking content and completion signals
    - Handles Gemini's parts-based content system (text, thought, functionCall)
    - Implements tool/function calling with proper format conversion
    - Maps between unified types and Gemini-specific formats (model vs assistant role)
    - Added test example matching other provider patterns
    - Fixed typo in AssistantMessage type (stopResaon -> stopReason) across all providers
  • refactor(ai): Add completion signal to onText/onThinking callbacks
    - Update LLMOptions interface to include completion boolean parameter
    - Modify all providers to signal when text/thinking blocks are complete
    - Update examples to handle the completion parameter
    - Move documentation files to docs/ directory
  • feat(ai): Add OpenAI Completions and Responses API providers
    - Implement OpenAICompletionsLLM for Chat Completions API with streaming
    - Implement OpenAIResponsesLLM for Responses API with reasoning support
    - Update types to use LLM/Context instead of AI/Request
    - Add support for reasoning tokens, tool calls, and streaming
    - Create test examples for both OpenAI providers
    - Update Anthropic provider to match new interface
  • feat(ai): Implement unified AI API with Anthropic provider
    - Define clean API with complete() method and callbacks for streaming
    - Add comprehensive type system for messages, tools, and usage
    - Implement AnthropicAI provider with full feature support:
      - Thinking/reasoning with signatures
      - Tool calling with parallel execution
      - Streaming via callbacks (onText, onThinking)
      - Proper error handling and stop reasons
      - Cache tracking for input/output tokens
    - Add working test/example demonstrating tool execution flow
    - Support for system prompts, temperature, max tokens
    - Proper message role types: user, assistant, toolResult