ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-05-11 00:20:19 +00:00

Author	SHA1	Message	Date
Iwan Kawrakow	a5b16b82bb	More gemv+add fusing	2025-10-26 09:32:02 +02:00
Iwan Kawrakow	7da6fb0979	Fuse Q, K, V gemv+add	2025-10-26 08:46:00 +02:00
Iwan Kawrakow	d0861f83f1	Fix TG fused up*nary(gate) when down cannot be fused The wrong memory buffer got used in that case	2025-10-26 08:38:25 +02:00
Iwan Kawrakow	8d01a40f5b	Disable assert	2025-10-25 15:21:13 +03:00
Iwan Kawrakow	5bd9cd3490	Add disagnostics	2025-10-25 11:12:16 +03:00
Iwan Kawrakow	a7d050426a	Somehow I forgot to change the ggml_type in the legacy template calls	2025-10-25 11:09:37 +03:00
Iwan Kawrakow	5ff69380db	Put iqk mmvq implementations into template instances	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	79d1c4ebb9	Split mmvq.cu and iqk_mmvq.cu into separate template instances	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	1fcae126cf	Also iqk quants	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	6b57074431	Fuse mul_mat_id and add_id into a single kernel for mmvq	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	4a08ac7241	Fusing mmvq also in non-MoE up+gate	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	196e73588c	Fusing also for iqk/trellis/repacked quants	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	3da71dcda2	Fused ffn_up*unary_op(ffn_gate) for MMVQ (with bias)	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	73c551aa9e	Fused ffn_up*unary_op(ffn_gate) for MMVQ (no bias) We see nearly 2% TG speedup for Ling-mini-2.0 and about 1% for DeepSeek-Lite.	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	b5cb6cd38e	WIP	2025-10-25 10:53:11 +03:00
Iwan Kawrakow	a46b5e337c	Args for MMVQ functions	2025-10-25 10:53:11 +03:00
Kawrakow	16f30fcf31	Change flash attention and fmoe to be on by default (#863 ) * Change fmoe to be on by default * Change default fmoe also in llama-bench * Change flash attention to be on by default --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-25 09:37:28 +03:00
Kawrakow	2522c97dc9	Faster tensor name formatting (#860 ) * Adding fused mul+multi_add + CPU implementation * fused mul+multi_add: command line argument to disable it * Faster tensor name formatting We gain ~1% for Ling-mini-2.0 when running on CUDA. --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-24 07:46:18 +03:00
Kawrakow	db3ba4999f	Fused mul + multi_add op (#858 ) * Adding fused mul+multi_add + CPU implementation * fused mul+multi_add: CUDA * fused mul+multi_add: command line argument to disable it --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-24 07:40:35 +03:00
Kawrakow	483cea527d	Fix experts mul node name (#857 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-23 09:46:01 +03:00
Kawrakow	0e1d33ca4a	Fuse add+add+fused_rms (#853 ) * Fuse add+add+fused_rms * Try this * Macro to easily enable/disable fusion * Various: * Check that all tensors involved are on the same device before applying fusion * Fuse sigmoid+scale+sum_rows+div * Fix the fused bailingmoe2 experts selection The issue there was that the bias was not per row, but per expert group, so only the first n_per_group biases were used for al experts. --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-22 16:18:11 +03:00
Kawrakow	8aa3c2ec5e	Hopefully this fixes #854 (#855 ) * Hopefully this fixes #854 * Also this one --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-21 19:07:23 +03:00
Kawrakow	caf9759c97	Fuse add + fused_rms_norm (CUDA) (#852 ) * Combine all calls to llm_build_norm to a single line so more easily check what kind of arguments are being passed by simply using grep. * Combine add + fused_rms_norm For many models this happens at each layer: the result of the layer is added to the ayer input, which then becomes the input to the next layer, which then is typically normalized via fused_rms_norm. --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-21 14:29:50 +03:00
Kawrakow	92231460cf	Fix fused grouped topk (#851 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-21 10:10:38 +03:00
Kawrakow	f5571e241e	cuda: use better block sizes for rms_norm (#845 ) * cuda: use better block sizes for rms_norm * Minor * Remove forgotten printf --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-21 08:12:48 +03:00
Kawrakow	5ae87f6cdf	Fix PR #842 (#844 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-20 11:35:57 +03:00
Kawrakow	22540cee60	Do not allocate KV cache for unused layers (#843 ) * Do not allocate KV cache for unused layers * Do not apply experts weight scale if it is 1 --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-20 10:09:39 +03:00
Kawrakow	1789de5994	Make ooae on by default and add to llama-bench (#842 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-20 08:32:41 +03:00
Kawrakow	2f5dae22e1	Change --n-cpu-moe to not keep expert biases on CPU (#841 ) * Change --n-cpu-moe to not keep expert biases ion CPU * Also for --cpu-moe --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-19 19:03:03 +03:00
Kawrakow	28d3e63805	Various fused ops around expert selection (#840 ) * Fuse sigmoid+add+grouped_topk+get_rows (CPU) * Fix CPU + CUDA but CUDA is somehow not 100% correct as I get a slightly different PPL (lower!) * Minor * Fuse sigmoid+add+topk+get_rows (CUDA) * Fuse sigmoid+add+topk+get_rows (CPU) * Fuse topk+view+get_rows+reshape+softmax (CPU) * Fuse topk+view+get_rows+reshape+softmax (CUDA) * cpu: turn off the openai topk fusing for now Something is not right and I don't see the bug. On the CPU one doesn't gain much if anything, so not a big loss. * Also fuse sum_rows and div --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-19 19:02:46 +03:00
Kawrakow	747f411da5	Grouped expert routing (CUDA) (#838 ) * WIP * cuda: grouped top_k * This is very slightly better --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-18 07:22:35 +03:00
ubergarm	2f951a8ab5	Ling-1T convert fixup (#837 ) * Conditionally write moe_shared_expert_intermediate_size Ling-1T config.json does not have `moe_shared_expert_intermediate_size`. Ling-flash-2.0a does have it. This small patch just makes the gguf_writer conditionally detect as needed. * Fix Ling-1T missing moe_shared_expert_intermediate_size Thanks CISC for the proper patch to include the needed values!	2025-10-17 07:52:31 +03:00
Kawrakow	dbfd151594	Grouped expert routing (CPU only) (#836 ) * Better argsort (CPU) * Attemt at grouped topk * This seems to do the trick for grouped experts routing * Cleanup * Trying to merge, something is not right * Working merged grouped top_k (CPU) * Add command line option to enable grouped expert routing * Add grouped expert routing option to llama-bench --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-16 14:57:02 +03:00
Kawrakow	ecf8f931ea	Better argsort (CPU) (#835 ) * Better argsort (CPU) * Minor --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-16 11:31:03 +03:00
Kawrakow	9d364b88ba	Adding Ling/Ring (a.k.a., Bailing-MoE2) support (#833 ) * Adding Ling/Ring (a.k.a., Bailing-MoE2) * Add expert group selection (not working, so turned off) * BailingMoE2 conversion * WIP * Bits and pieces --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-15 14:20:40 +03:00
Kawrakow	8d0d01a593	gpt-oss: duplicate experts biases when necessary (#829 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-14 14:38:40 +03:00
Viktor Ivakin	41bdd86555	Fix incomplete utf-8 characters in streaming text completions (#810 )	2025-10-13 16:25:29 +03:00
Kawrakow	9724ea9213	Attention mask tweaks for better long context performance (#825 ) * Parallelize mask We see non-negligible PP gains for long contexts. More importantly, the strange drop in performance observed for GPT-OSS for context >= 32k tokens is gone. * Whith FA on, create mask as f16 directly * WIP * Reduce KQ mask padding to 16 Why was it 64 in the first place? I don't observe any issues, while TG performance for long contexts improves by 2-4%. --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-13 14:01:11 +03:00
Kawrakow	1db0c490be	Fix PATH_MAX not defined on Windows (#828 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-13 09:25:57 +03:00
Kawrakow	0030bc89c9	Fix performance regression introduced in #823 (#826 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-13 08:09:55 +03:00
Kawrakow	0ad1d34090	Enable and clean up compiler warnings in src (#824 ) * WIP: enable and clean up warnings in src * All warnings handled --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-11 16:01:13 +03:00
Kawrakow	335a1f9b71	Refactor file llama.cpp (#823 ) * llama_model and llama_hparams * llama_build_context Surprisingly small reduction in llama.cpp compile time given the reduction in LOCs (22k -> 14k) * LLM_TN llama.cpp compilation: 50 s -> 33 s * llama_quantize * arch names * All graph building is now in llm-build-context.cpp * hparams loading llama.cpp is now just 9300 LOC, but still takes 32 seconds to compile. * We are now at 6 seconds to build the src folder * load -> create We are not actually loading the tensors, but just creating them. --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-10-11 11:35:20 +03:00
AesSedai	f649e36a61	Remove duplicate 99% KLD output, add additional percentiles to match mainline (#817 )	2025-10-05 07:13:32 +02:00
Downtown-Case	6051ba25ee	Mark some multi-prediction tensors as not required. (#814 )	2025-10-01 20:37:31 +02:00
Kawrakow	e94d1a92a5	Attempt to fix AVX2 FA (#807 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-09-30 08:06:53 +02:00
Kawrakow	3d4977cb6e	Fix gemma3 vision (#803 ) * Remove unnecessary assert in im2col * Remove unnecessary assert in im2col (CPU) --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-09-27 11:15:32 +02:00
Kawrakow	95780cddc9	Move minja and nlohmann/json to vendor (#802 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-09-27 09:12:35 +02:00
Kawrakow	5064ff8a54	Remove stb_image.h copy in common - it is now in vendor (#801 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-09-27 08:55:42 +02:00
Kawrakow	87e4762720	Port mdmd from mainline + Qwen2/2.5-VL support (#798 ) * Add mtmd: the beginning * Add mtmd: mtmd.cpp compiles * Add mtmd: clip initialization compiles * Add mtmd: clip.cpp compiles * Add mtmd: builds successfully * Add CPU implementation for GGML_OP_GLU * Add CUDA implementation for GGML_OP_GLU * Add CPU implementation for GGML_OP_CONV_2D and GGML_OP_CONV_2D_DW * Add CUDA implementation for GGML_OP_CONV_2D and GGML_OP_CONV_2D_DW * Add mtmd: refresh CPU rope * Add mtmd: refresh CUDA rope * Add mtmd: add Qwen2-VL * Add mtmd: Qwen2.5-VL text seems to work with this change * Add mtmd: fix swiglu * Add mtmd: use LOG_TEE so generated tokens show up in terminal * Add mtmd: do not attempt to load a GPU backend if none are available * GLU, not GPU * Fix typo * Fix new/free mismatch * LOG stuff * Add mtmd: this fixes gibberish on second image --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2025-09-27 08:45:29 +02:00
firecoperana	367654f99e	sync: vendor (#799 ) Co-authored-by: firecoperana <firecoperana>	2025-09-26 18:22:47 +02:00

1 2 3 4 5 ...

3944 Commits