test (closed) - #1
Closed
linsen458-spec wants to merge 65 commits into
Closed
linsen458-spec wants to merge 65 commits into
linsen458-spec wants to merge 65 commits into
Conversation
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
…er what is compiled (ggml-org#28079) * CUDA: add configurable FA quant combinations Assisted-by: Codex * remove all flags but , add runtime fallback with warning for uncompiled combination * Update docs/build.md Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * apply code review comments --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
…gml-org#28552) Recreated from ggml-org#24546 --------- Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz> * CUDA: pick MMQ tile size against ncols_opt set on the host side Assisted-by: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_011SYPfRhKoUpU3gMsGxq6go --------- Co-authored-by: ravel7524 <58877666+ravel7524@users.noreply.github.com> Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz>
…8631) Contributes to ggml-org#4574 Co-authored-by: linsen <linsen@insta360.com>
* Revert "py : bump numpy to 2.4.6 (ggml-org#28649)" This reverts commit 9cf3bf2. * bump numpy to 2.2.6
…g maxComputeWorkGroupCount (ggml-org#28592) * divide workload to 2D This is to workaround FILL exceeding maxComputeWorkGroupCount for Intel GPUs on Qwen 3.8 flash next * minor change * Fixed comment
* hexagon: vectorize RoPE theta cache on v75 * hexagon: vectorize MROPE/IMROPE theta pick * hexagon: tighten NEOX RoPE rotate and aligned tail copy * hex-rope: use inplace rope for all scenarios * hex-rope: remove ctx->spad usage and legacy timers * hex-rope: add kernel params and enforce vtcm reqs at the host * hex-rope: cleanup unused params and tighten the mode checks * hex-rope: add missing ops header --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
…ml-org#25773) * vulkan: use spec constant for mul mat type_a vulkan: use map for mul_mm shapes cleanup fix indentation fix cm2 and shmem init fix cm2 spec constants fix cm2 bindings consolidate shmem tables and reduce size by type spec constant fix compiler warning fix missing Q2_0 type fix unused warning when integer dot glslc support is missing use minimal shmem size 8 instead of 1 to workaround cm2 compiler bug fix missing Q2_0 type in cm2 matmul fix types * remove LUT quants from unified shader * clean up * restore coopmat2 q4_k/q5_k optimization * split out q4_k/q5_k cm2 shader to fix Ampere regression * revert iq shmem table renames * simplify cm2 code with single uint8_t buffer * fix fp4 extension use switch being overwritten by generic shader * clean up * adapt TQ1_0 changes * adapt ggml-org#27471 f16 Intel tuning changes
* model: fix all granite family parameter counts Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * model: fix additional include, add missing `A` prefix for active experts Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * model: fix code alignment, rm unused 40 block case Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> --------- Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* ggml-cpu: add `ggml_vec_dot_q1_0_q8_0` support Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ggml-cpu: clean up variable naming for understanding Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * docs: update support for Q1_0 Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> --------- Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
ggml-cpu: clean comments Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
…28101) * vulkan : add command-buffer debug labels for GPU profilers Co-authored-by: gabby-zy <z2262718160@gmail.com> Assisted-by: Claude Code * vulkan : close the queue debug label with the label struct --------- Co-authored-by: gabby-zy <z2262718160@gmail.com>
…rg#28330) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
* override function for n_h_l * narrow change for extracting nested attribute * simpler change; combines has_moe_params * Apply suggestion from @CISC Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…l-org#28341) The Imagination proprietary Vulkan compiler returns VK_ERROR_UNKNOWN from vkCreateComputePipelines for every dequant mul_mat_vec shader built with the subgroup-only reduction that requires a subgroup size >= 16. That covers the k-quants, the i-quants, TQ2_0, MXFP4 and NVFP4. ggml rethrows, so the first generated token of any such model kills the process. Reproduced on a Pixel 11 Pro (PowerVR C-Series CXTP-48-1536 MC1, driver 1.662.3024, subgroup size 128, min 32, max 128). The failure is independent of subgroup size: 32, 64 and 128 all fail, as does dropping the full-subgroups flag and the required-subgroup-size pNext. The legacy quants, which use the plain subgroup reduction, compile and run fine. The shared-memory reduction variant compiles and matches the CPU reference for q2_K, q3_K, q4_K, q5_K and q6_K. The hybrid variant also compiles but costs 27% of token throughput (3.78 vs 5.20 t/s on Qwen3.5-2B-Q4_K_M).
…rg#28700) This commit updates the version parsing in make-release-checks.sh to use sed instead of grep. The motivation for this is that currently when running this script on macos it errors: ```console $ ./scripts/make-release-checks.sh --dry-run grep: invalid option -- P usage: grep [-abcdDEFGHhIiJLlMmnOopqRSsUVvwXxZz] [-A num] [-B num] [-C[num]] [-e pattern] [-f file] [--binary-files=value] [--color=when] [--context[=num]] [--directories=action] [--label] [--line-buffered] [--null] [pattern] [file ...] ``` With the changes in this commit it is possible to run this without failure.
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation
…gml-org#28687) - Move Windows ARM64 CUDA 13.4 builds from the Developer Preview archives to the 13.4.1 GA redistributables
* vulkan: optimize m=1 mul_mat by swapping A/B * vulkan: Improve small M perf Allow split_k with small M. Make small vs med tile selection (for coopmat2) depend on M, not just N.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…org#28816) Clang stores the modification time of the precompiled header sources inside the header and refuses the header when they differ. A cached header restored from another checkout carries the timestamps of that checkout, so the build fails. The option covers the compilers ccache treats as MSVC while they are clang underneath, clang-cl and the Intel LLVM drivers.
…emas (ggml-org#28736) * common : implement common_schema types * common : implement a json schema optimizer * common : reduce optimizations * common : refactor json-schema-to-grammar to use common_schema * common : use common_trie * common/schema : implement type/kind resolution * cont : cleanup * cont : remove common_chat_tool_parameters * cont : simplify schema resolution * cont : pass common_schema through the json-schema-to-grammar builder * cont : cleanup * cont : move enums under common_schema and add type enum * cont : reduce test cases * cont : clean up * cont : clean up * refactor : rename common_schema_parse to common_schema_from_json * tests : fix gcc dangling-reference warning in test-json-schema * tests : take the schema label as const char * to satisfy gcc dangling-reference * refactor : rename common_schema_builder parse_* methods to build_* * cont : fix may_be_string * cont : properly handle empty tool parameters * cont : add tests for empty $ref * cont : remove dead code * cont : update docs * cont : make "{}" mean any object for json_object as well * cont : restore (min|max)Length to imply string type * cont : rename common_schema to common_chat_schema
* add LOG_JSON macro * fit: add demo LOG_JSON
* chat : improve schema support in qwen3 parser * cont : clean up grammar a bit
unset_reserved_args() dropped LLAMA_API_KEY from the base preset before it was merged into the per-model presets, so children spawned by the router only re-validated against --api-key-file keys: a client that authenticated with the router's --api-key key could list models but got 401 from every chat completion (ggml-org#28820). Keep LLAMA_API_KEY in the preset so children accept the same key union the router accepts (they bind to router-assigned loopback ports); the /api/models preset rendering masks it separately. Fixes ggml-org#28820
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.