ollama37

mirror of https://github.com/dogkeeper886/ollama37.git synced 2025-12-20 20:57:01 +00:00

Author	SHA1	Message	Date
Shang Chieh Tseng	82ab6cc96e	Refactor model unload: each test cleans up its own model - TC-INFERENCE-003: Add unload step for gemma3:4b at end - TC-INFERENCE-004: Remove redundant 4b unload at start - TC-INFERENCE-005: Remove redundant 12b unload at start Each model size test now handles its own VRAM cleanup. Workflow-level unload remains as safety fallback for failures. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2025-12-17 17:20:44 +08:00
Shang Chieh Tseng	806232d95f	Add multi-model inference tests for gemma3 12b and 27b - TC-INFERENCE-004: gemma3:12b single GPU test - TC-INFERENCE-005: gemma3:27b dual-GPU test (K80 layer split) - Each test unloads previous model before loading next - Workflows unload all 3 model sizes after inference suite - 27b test verifies both GPUs have memory allocated	2025-12-17 17:01:25 +08:00
Shang Chieh Tseng	1a185f7926	Add comprehensive Ollama log checking and configurable LLM judge mode Test case enhancements: - TC-RUNTIME-001: Add startup log error checking (CUDA, CUBLAS, CPU fallback) - TC-RUNTIME-002: Add GPU detection verification, CUDA init checks, error detection - TC-RUNTIME-003: Add server listening verification, runtime error checks - TC-INFERENCE-001: Add model loading logs, layer offload verification - TC-INFERENCE-002: Add inference error checking (CUBLAS/CUDA errors) - TC-INFERENCE-003: Add API request log verification, response time display Workflow enhancements: - Add judge_mode input (simple/llm/dual) to all workflows - Add judge_model input to specify LLM model for judging - Configurable via GitHub Actions UI without code changes 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2025-12-16 23:27:57 +08:00
Shang Chieh Tseng	ebcca9f483	Add model warmup step to TC-INFERENCE-001 Tesla K80 needs ~60-180s to load model into VRAM on first inference. Add warmup step with 5-minute timeout to preload model before subsequent inference tests run. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2025-12-15 21:38:09 +08:00
Shang Chieh Tseng	3f3f68f08d	Remove TC-INFERENCE-004: CUBLAS Fallback Verification Redundant test - if TC-INFERENCE-002 (Basic Inference) passes, CUBLAS fallback is already working. Any errors would cause inference to fail, making a separate error-check test unnecessary. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2025-12-15 20:38:35 +08:00
Shang Chieh Tseng	d11140c016	Add GitHub Actions CI/CD pipeline and test framework - Add .github/workflows/build-test.yml for automated testing - Add tests/ directory with TypeScript test runner - Add docs/CICD.md documentation - Remove .gitlab-ci.yml (migrated to GitHub Actions) - Update .gitignore for test artifacts 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2025-12-15 14:06:44 +08:00

6 Commits