Category detail

Local LLM / Open Modelspage 5/8

Local LLM / Open Models カテゴリの更新の続きです。ページを移動してもカテゴリ文脈と主要指標を維持します。Continue browsing Local LLM / Open Models updates while keeping category context and key metrics in view.

Total230現在のカテゴリ一覧Current category listing
Showing30このページの表示件数Entries on this page
Last 7d48直近7日の掲載数Entries in the latest 7 days
Vs prev 7d+129%その前の7日間と比較Compared with the previous 7 days
Page5/8静的ページ位置Static page position

All articlespage 5/8 · 230 total

新着順Newest first
Sun, Jul 261 entries
公式OfficialLocal Models·Ollama Releases

Ollama v0.32.4 リリースOllama Releases v0.32.4

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollama v0.32.4がリリースされ、ローカルLLM実行環境の安定性と品質が改善された。小規模なパッチリリースだが、継続的なメンテナンスの一環として重要。

AI SUMMARYOllama v0.32.4 is a patch release delivering bug fixes and stability improvements to the local LLM runtime, keeping the platform reliable for self-hosted AI workloads.

Ollama Releases v0.32.4media
Sat, Jul 254 entries
コミュニティCommunityLocal Models·Zenn LLM

非力なGPUでローカルLLMは動くか――Gemma 4 E2B QATの実験環境とPythonコードを公開A developer shares a reproducible experiment running Gemma 4 E2B QAT on a…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約低スペックGPU環境でGemma 4 E2B QATモデルを動作させる実験を行い、その環境構成と再現可能なPythonコードを公開。手頃なハードウェアでローカルLLMを活用できる可能性を示した。

AI SUMMARYA developer shares a reproducible experiment running Gemma 4 E2B QAT on a low-end GPU, publishing the full environment setup and Python code to help others run local LLMs on modest hardware.

非力なGPUでローカルLLMは動くか――Gemma 4 E2B QATの実験環境とPythonコードを公開og
コミュニティCommunityLocal Models·Zenn LLM

Gemma 4 12BをiPhoneで投機デコードする:2.4倍高速化とA19最適化This article details how speculative decoding applied to Gemma 4 12B on Apple's…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約iPhoneのA19チップ上でGemma 4 12Bを動作させる際に投機的デコードを適用し、推論速度を最大2.4倍改善した手法を解説した記事。エッジデバイスでの大規模モデル実用化に向けた具体的な最適化アプローチとして注目される。

AI SUMMARYThis article details how speculative decoding applied to Gemma 4 12B on Apple's A19 chip achieves up to 2.4× inference speedup on iPhone, offering practical techniques for running large models efficiently on edge devices.

公式OfficialLocal Models·Ollama Releases

Ollama v0.32.4-rc0 リリースv0.32.4-rc0: model: add Laguna MLX support (#17237)

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollama v0.32.4がリリースされ、ローカルLLM実行環境の安定性と品質が改善された。ユーザーは最新版へのアップデートが推奨される。

AI SUMMARYOllama v0.32.4 is a patch release delivering bug fixes and stability improvements to the local LLM runtime, keeping the toolchain current for self-hosted AI workflows.

v0.32.4-rc0: model: add Laguna MLX support (#17237)media
コミュニティCommunityLocal Models·Zenn LLM

Gemma 4 12B を Core ML で 128K コンテキストで動かすThis article explains how to run Gemma 4 12B with a 128K context window on…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Apple Silicon 上で Core ML を使い Gemma 4 12B を 128K トークンのコンテキスト長で動作させる手順を解説した記事で、ローカル環境での大規模モデル活用の可能性を示している。

AI SUMMARYThis article explains how to run Gemma 4 12B with a 128K context window on Apple Silicon using Core ML, demonstrating that large open models can be deployed locally without cloud infrastructure.

Fri, Jul 242 entries
公式OfficialLocal Models·Ollama Releases

Ollama v0.32.2 リリースOllama Releases v0.32.2

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollama v0.32.2がリリースされ、ローカルLLM実行環境の安定性と品質が改善された。継続的なメンテナンスにより信頼性が向上している。

AI SUMMARYOllama v0.32.2 is a patch release delivering bug fixes and stability improvements to the local LLM runtime, keeping deployments reliable and up to date.

Ollama Releases v0.32.2media
公式OfficialLocal Models·Ollama Releases

Ollama v0.32.3 リリースOllama Releases v0.32.3

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollama v0.32.3がリリースされ、ローカルLLM実行環境の安定性と品質が改善された。ユーザーは最新版へのアップデートが推奨される。

AI SUMMARYOllama v0.32.3 is a patch release delivering bug fixes and stability improvements to the local LLM runtime, keeping the platform reliable for self-hosted deployments.

Ollama Releases v0.32.3media
Thu, Jul 236 entries
コミュニティCommunityLocal Models·Zenn LLM

VRAMに乗らないMoEをNVMe+GPU推論で動かす:Hypura/llama.cpp/TurboQuant解説This article explains how to run large MoE models that exceed VRAM capacity by…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約VRAMに収まらない大規模MoEモデルをNVMeストレージとGPUを組み合わせて実用的に推論する手法を、Hypura・llama.cpp・TurboQuantの三ツールを軸に解説した記事。コンシューマー環境でも巨大モデルを動かせる可能性を示す点で注目に値する。

AI SUMMARYThis article explains how to run large MoE models that exceed VRAM capacity by offloading layers to NVMe storage while leveraging GPU acceleration, using Hypura, llama.cpp, and TurboQuant. It matters because it opens a practical path for running frontier-scale models on consumer hardware.

公式OfficialLocal Models·Ollama Releases

v0.32.3-rc0: Lagunaモデルをupstream llama.cppに同期v0.32.3-rc0: model: align Laguna with upstream llama.cpp (#17335)

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollamaのリリース候補v0.32.3-rc0では、LagunaモデルのアーキテクチャをアップストリームのLlama.cppの実装に合わせる修正が行われた。互換性と動作精度の向上が目的。

AI SUMMARYRelease candidate v0.32.3-rc0 aligns Ollama's Laguna model implementation with upstream llama.cpp, ensuring compatibility and correctness with the reference architecture.

v0.32.3-rc0: model: align Laguna with upstream llama.cpp (#17335)media
コミュニティCommunityLocal Models·Zenn LLM

ollama の入力切り捨てをレスポンスだけで検知する — 3回作り直した記録A practical account of detecting silent input truncation in ollama—where…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約ollama がコンテキスト長を超えた入力を無警告で切り捨てる問題に対し、レスポンスのみを手がかりに切り捨てを検知する手法を3度の試行錯誤を経て確立した実践記録。ローカルLLM運用の信頼性向上に役立つ知見をまとめている。

AI SUMMARYA practical account of detecting silent input truncation in ollama—where prompts exceeding the context window are cut without warning—using only the model response as a signal, refined through three redesigns. The findings help improve reliability when running LLMs locally.

コミュニティCommunityLocal Models·Zenn LLM

Voicebox に Jetson 8GB の Bonsai 27B を繋いだ話 — OpenAI互換APIをRustで全部書いた理由The author ran Bonsai 27B on a Jetson with 8 GB RAM and built a full…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約8GBメモリのJetson上でBonsai 27Bを動かし、VoiceboxからOpenAI互換APIで呼び出せるよう、RustでAPIサーバーをゼロから実装した経緯と技術的な選択理由を解説している。

AI SUMMARYThe author ran Bonsai 27B on a Jetson with 8 GB RAM and built a full OpenAI-compatible API server in Rust to connect it to Voicebox, explaining why Rust was chosen over existing solutions.

コミュニティCommunityLocal Models·Simon Willison's Weblog

Thomas Ptacek の言葉を引用Quoting Thomas Ptacek

重要度 InfoInformational深掘り候補 · 技術記事 · Local LLM / Open ModelsDeep-dive candidate · technical post · Local LLM / Open Models

AI要約セキュリティ研究者 Thomas Ptacek によるローカル LLM に関する見解を Simon Willison が取り上げ、その実用性や限界について注目すべき視点を紹介している。

AI SUMMARYSimon Willison highlights a notable take from security researcher Thomas Ptacek on local LLMs, surfacing an expert perspective worth attention in the ongoing conversation about their practical value.

コミュニティCommunityLocal Models·Simon Willison's Weblog

AIラボは「pelicanmaxxing」をしているのか?Are AI labs pelicanmaxxing?

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Simon Willison氏が、AIラボがベンチマーク向けに過剰最適化する「pelicanmaxxing」と呼ばれる現象を考察し、実用性より指標改善を優先するリスクを指摘した記事。

AI SUMMARYSimon Willison examines whether AI labs are "pelicanmaxxing" — over-optimizing for benchmarks and metrics at the expense of genuine usefulness, and what that means for the field.

Are AI labs pelicanmaxxing?media
Wed, Jul 223 entries
公式OfficialLocal Models·Ollama Releases

v0.32.2-rc3: 統合テストのエントリーポイントを刷新 (#16560)v0.32.2-rc3: test: revamp integration test entrpoints (#16560)

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollamaのリリース候補v0.32.2-rc3では、統合テストのエントリーポイントが整理・刷新され、テスト基盤の品質と保守性が向上した。

AI SUMMARYOllama v0.32.2-rc3 revamps integration test entry points, improving test infrastructure maintainability ahead of the stable release.

v0.32.2-rc3: test: revamp integration test entrpoints (#16560)media
公式OfficialLocal Models·Ollama Releases

Ollama v0.32.2-rc2 リリースv0.32.2-rc2: CI: fix missing CUDA v13.4 sub-package (#17288)

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollama v0.32.2がリリースされ、ローカルLLM実行環境の安定性と品質が改善された。ユーザーは最新版へのアップデートが推奨される。

AI SUMMARYOllama v0.32.2 is a patch release bringing stability improvements and bug fixes to the local LLM runtime, keeping the platform reliable for self-hosted AI workloads.

v0.32.2-rc2: CI: fix missing CUDA v13.4 sub-package (#17288)media
公式OfficialLocal Models·Ollama Releases

v0.32.2-rc1: サーバーが最初のバイト受信前にダウンロードの停止を検出する機能を追加v0.32.2-rc1: server: detect download stalls before the first byte (#17259)

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollamaのv0.32.2-rc1では、モデルダウンロード中に最初のバイトが届く前にスタックを検出できるよう改善され、ダウンロード失敗時の検知が早くなった。

AI SUMMARYOllama v0.32.2-rc1 improves download reliability by detecting stalls before the first byte arrives, allowing the server to catch hung downloads earlier than before.

v0.32.2-rc1: server: detect download stalls before the first byte (#17259)media
Tue, Jul 217 entries
コミュニティCommunityLocal Models·Simon Willison's Weblog

Nativ: Mac でAIモデルをローカル実行するアプリNativ: Run AI models locally on your Mac

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約NativはMac上でAIモデルをローカル動作させるアプリで、プライバシーを保ちながらオフラインでLLMを活用できる点が注目される。

AI SUMMARYNativ is a macOS app that lets users run AI models locally, enabling private and offline LLM usage without relying on cloud services.

コミュニティCommunityLocal Models·Qiita LLM

Claude Fable 5 を9Bモデルに蒸留? 100万トークンの超長文推理モデル「Qwythos-9B」を4GBのVRAMで動かすQwythos-9B is a purported Claude Fable 5 distillation that supports 1M-token…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Qwythos-9BはClaude Fable 5からの蒸留とされる9Bパラメータの推論モデルで、100万トークンのコンテキストを持ちながら4GB VRAMで動作する点が注目される。

AI SUMMARYQwythos-9B is a purported Claude Fable 5 distillation that supports 1M-token context while running on just 4 GB of VRAM, making long-context reasoning accessible on consumer hardware.

コミュニティCommunityLocal Models·Qiita LLM

小さなLLM(Llama-3.2-1B)をQLoRAでファインチューニングしてFunction Callingを覚えさせてみたThis article demonstrates fine-tuning the compact Llama-3.2-1B model with QLoRA…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Llama-3.2-1BというコンパクトなモデルにQLoRAを用いてFunction Callingを学習させる手法を解説しており、限られたリソースでもツール呼び出し能力を獲得できることを示している。

AI SUMMARYThis article demonstrates fine-tuning the compact Llama-3.2-1B model with QLoRA to enable function calling, showing that tool-use capabilities can be taught even on limited hardware.

コミュニティCommunityLocal Models·Zenn LLM

ollama は長い入力の中間部分を無音で切り捨てる — 実測で半分しか処理されない問題Benchmarking reveals that ollama silently drops the middle portion of inputs…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約ollama はデフォルトのコンテキスト長を超えた入力を受け取ると、警告なしに中間部分を削除することが実測で判明した。ユーザーが気づかないまま重要な情報が欠落するため、num_ctx の明示的な設定が必要となる。

AI SUMMARYBenchmarking reveals that ollama silently drops the middle portion of inputs exceeding the default context length, retaining only about half the content without any warning. This silent truncation can cause critical information loss, making explicit num_ctx configuration essential.

コミュニティCommunityLocal Models·Zenn LLM

AIの評価を報酬にする強化学習は何をしているのか — GRPOの1ステップを数字で追うThis article walks through a single GRPO optimization step with concrete…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約GRPOアルゴリズムの1ステップを具体的な数値で丁寧に解説し、AIの出力評価を報酬信号として用いる強化学習の仕組みを直感的に理解できるようにした記事。実装の背景を理解したい実践者にとって有益な内容。

AI SUMMARYThis article walks through a single GRPO optimization step with concrete numbers, demystifying how AI-generated evaluations are used as reward signals in reinforcement learning for language models.

公式OfficialLocal Models·Ollama Releases

v0.32.2-rc0: CUDA v12 の Linux 向けに CC 10.0 サポートを追加v0.32.2-rc0: cuda: add CC 10.0 for linux in CUDA v12 (#17025)

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollama v0.32.2-rc0 では、Linux 環境の CUDA v12 に Compute Capability 10.0 対応が追加され、最新世代の NVIDIA GPU でのローカル LLM 実行が可能になります。

AI SUMMARYOllama v0.32.2-rc0 adds Compute Capability 10.0 support for Linux under CUDA v12, enabling local LLM inference on the latest generation of NVIDIA GPUs.

v0.32.2-rc0: cuda: add CC 10.0 for linux in CUDA v12 (#17025)media
コミュニティCommunityLocal Models·Simon Willison's Weblog

中国製AIモデルを恐れる必要はあるか?Who’s Afraid of Chinese Models?

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約中国製LLMの利用に対するセキュリティや政治的懸念を検討し、ローカル実行の文脈でそのリスクと実用性を評価した考察記事。開発者がどう向き合うべきかを論じている。

AI SUMMARYSimon Willison examines the fears and practical realities around using Chinese-origin LLMs, weighing security and political concerns against their performance, especially in local deployment scenarios.

Who’s Afraid of Chinese Models?media
Mon, Jul 203 entries
コミュニティCommunityLocal Models·Qiita LLM

QSpec の論文要点整理A structured breakdown of the QSpec paper, explaining its core ideas around…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約QSpec に関する論文の主要なポイントをまとめ、量子化仕様の設計思想と実用上の意義を解説した記事。ローカル LLM の量子化運用に関心を持つ実践者にとって有益な参照資料となる。

AI SUMMARYA structured breakdown of the QSpec paper, explaining its core ideas around quantization specification design and why it matters for practical local LLM deployment.

コミュニティCommunityLocal Models·Zenn LLM

BIRD:ブートストラップ自己蒸留で推論CoTを64%圧縮しつつ精度も向上BIRD is a bootstrap self-distillation method that compresses chain-of-thought…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約BIRDはモデル自身の推論チェーンをブートストラップ自己蒸留で圧縮する手法で、CoTトークン数を最大64%削減しながら精度を維持・向上させる。推論コストの削減と性能の両立を示した点で注目に値する。

AI SUMMARYBIRD is a bootstrap self-distillation method that compresses chain-of-thought reasoning traces by up to 64% while maintaining or improving accuracy. This matters because it offers a practical path to reducing inference costs without sacrificing model performance.

コミュニティCommunityLocal Models·Qiita LLM

18社・88モデルを1つのAPIキーで利用できる「AICraft」が公開——ルーティングが自動で最適モデルを選択AICraft is a newly released service that unifies 88 models from 18 providers…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約AICraftは18社・88種類のLLMを単一APIキーで利用でき、リクエスト内容に応じて最適なモデルを自動選択するルーティング機能を備えたサービス。複数プロバイダーの管理コストを削減しつつ、常に適切なモデルを活用できる点が特徴です。

AI SUMMARYAICraft is a newly released service that unifies 88 models from 18 providers under a single API key, automatically routing each request to the most suitable model. This simplifies multi-provider LLM integration and removes the overhead of managing separate credentials and model selection logic.

18社・88モデルを1つのAPIキーで。ルーティングが勝手に最適モデルを選んでくれる「AICraft」を公開しましたog
Sun, Jul 194 entries
コミュニティCommunityLocal Models·Zenn LLM

ローカルLLM study3: gemma4:e2b vs Ornith-1.0-9B vs qwen3:14bを徹底比較するThis article benchmarks three locally-runnable LLMs—gemma4:e2b, Ornith-1.0-9B,…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約ローカル環境で動作する3つのLLM(gemma4:e2b、Ornith-1.0-9B、qwen3:14b)を複数のタスクで比較検証し、それぞれの性能差と用途適性を明らかにしている。

AI SUMMARYThis article benchmarks three locally-runnable LLMs—gemma4:e2b, Ornith-1.0-9B, and qwen3:14b—across multiple tasks to highlight their performance differences and practical use cases.

ローカルLLM study3: gemma4:e2b vs Ornith-1.0-9B vs qwen3:14bを徹底比較するog
コミュニティCommunityLocal Models·Zenn LLM

1-bit LLM「Bonsai」活用ガイド — 1.15GB で動く 8B モデルをローカルで使い倒すThis guide covers how to run Bonsai, a 1-bit quantized 8B LLM that fits in just…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約わずか1.15GBのメモリで動作する8Bパラメータの1-bit LLM「Bonsai」をローカル環境で活用する方法を解説。超軽量ながら実用的な推論が可能で、一般的なPCでも大規模モデルを手軽に運用できる点が注目される。

AI SUMMARYThis guide covers how to run Bonsai, a 1-bit quantized 8B LLM that fits in just 1.15 GB, entirely on local hardware. Its extreme compression makes powerful language models accessible on everyday consumer machines without cloud dependency.

コミュニティCommunityLocal Models·Zenn LLM

ローカルLLM(Ollama)にJSONを厳密に返させる — 自分専用ニュースbot開発記 #1This article explains how to enforce strict JSON output from a local LLM…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約OllamaでローカルLLMを動かす際に、structured outputを使ってJSONスキーマに厳密に準拠したレスポンスを得る方法を解説した開発記録。プロンプト工夫だけでは不安定だった出力を安定させる実践的な知見を共有している。

AI SUMMARYThis article explains how to enforce strict JSON output from a local LLM running on Ollama using structured output schemas, solving the instability that comes from prompt-engineering alone. It serves as the first entry in a series building a personal news bot.

コミュニティCommunityLocal Models·Zenn LLM

RAGFlowが日本語を中国語に変換する問題を回避するため、LlamaIndexで日英RAGを自作した話Faced with RAGFlow incorrectly converting Japanese text to Chinese, the author…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約RAGFlowが日本語テキストを誤って中国語に変換してしまう不具合を受け、著者がLlamaIndexを使って日本語・英語対応のRAGシステムをスクラッチで構築した経緯と実装方法を紹介している。

AI SUMMARYFaced with RAGFlow incorrectly converting Japanese text to Chinese, the author built a custom bilingual Japanese-English RAG pipeline from scratch using LlamaIndex, sharing the implementation details and lessons learned.