oh-my-openagent

Author	SHA1	Message	Date
YeonGyu-Kim	41c7c71d0d	Remove unused benchmark OpenAI SDK dependency	2026-03-11 15:33:05 +09:00
minpeter	8fb5949ac6	fix(benchmarks): address review feedback on error handling and validation - headless.ts: emit error field on tool_result when output starts with Error: - test-multi-model.ts: errored/timed-out models now shown as RED and exit(1) - test-multi-model.ts: validate --timeout arg (reject NaN/negative) - test-edge-cases.ts: use exact match instead of trim() for whitespace test - test-edge-cases.ts: skip file pre-creation for create-via-append test Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>	2026-02-27 01:44:51 +09:00
minpeter	04f50bac1f	feat(benchmarks): add hashline-edit test suites (46 tests) Ported from code-editing-agent benchmark: - test-edit-ops.ts: 21 basic edit operations (replace, append, prepend, delete, batch, range) - test-edge-cases.ts: 25 edge cases (unicode, long lines, whitespace, special chars, file creation) - test-multi-model.ts: multi-model comparison runner Verified 21/21 + 25/25 (100%) with Minimax M2.5 via FriendliAI. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>	2026-02-27 01:37:49 +09:00
minpeter	d1a0a66dde	feat(benchmarks): add hashline-edit benchmark agent and deps Standalone headless agent using Vercel AI SDK v6 with FriendliAI provider. Imports hashline-edit pure functions directly from src/ for benchmarking the edit tool against LLMs (Minimax M2.5 via FriendliAI). Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>	2026-02-27 01:37:40 +09:00

4 Commits