ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-05-11 16:40:16 +00:00

Author	SHA1	Message	Date
Brian	f0f87a2f2b	Added a single test function script and fix debug-test.sh to be more robust (#7279 ) * run-single-test.sh: added a single test function script and fix debug-test.sh to be more robust * debug-test.sh: combined execute and gdb test mode via -g flag * debug-test.sh: refactor * debug-test: refactor for clarity * debug-test.sh: comment style changes * debug-test.sh: fix gdb	2024-05-17 22:40:14 +10:00
Aarni Koskela	5ddcb26ba6	py : convert-hf-to-gguf-update improvements (#7340 ) * convert-hf-to-gguf-update: automate updating * convert-hf-to-gguf-update: improve download * share requests session for performance * create directories only when needed, don't skip downloads when empty directory encountered * be more graceful about errors	2024-05-17 15:11:45 +03:00
fairydreaming	16472b59b2	llama : use n_embd_head_v when reshaping kqv (#7327 ) * llama : use n_embd_head_v instead of n_embd_head_k when reshaping kqv * llama : use n_embd_v_gqa and n_embd_head_v instead of n_embd_k_gqa and n_embd_head_k when making a view of cached value vectors. --------- Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>	2024-05-17 14:24:38 +03:00
Johannes Gäßler	228b8cd135	tokenization: add warning for double BOS (#7332 )	2024-05-17 09:59:57 +02:00
Herman Semenov	6b3ce6e8b3	ggml-quants, llama : removed excess checks (#7274 )	2024-05-17 10:08:49 +03:00
amd-lalithnc	938cbd3d10	convert : fix Qwen/Qwen-7b conversion (#7308 )	2024-05-17 10:01:58 +03:00
Radoslav Gerganov	d0b0e30ad0	server : add support for the RPC backend (#7305 ) ref: #7292	2024-05-17 10:00:17 +03:00
Justine Tunney	3afea04b45	ggml : rewrite silu and softmax for cpu (#7154 ) This change upstreams llamafile's vectorized expf() functions. This lets us compute softmax and silu more accurately than the short[65536] lookup table that GGML previously used to make this operation go faster. We can support aarch64 and sse2+ with the worst case rounding error of 2ulp. It makes make -j8 tests && ./tests/test-backend-ops -o SOFT_MAX -b CPU perf go 1.5x faster for SSE2+FMA, 1.9x faster for AVX2+FMA and 2.1x on AVX512	2024-05-17 09:58:52 +03:00
Leon Knauer	7762ef55ba	[Server] Added --verbose option to README [no ci] (#7335 )	2024-05-17 10:11:03 +10:00
Pierrick Hymbert	99d7a99cef	Revert "server bench: fix bench not waiting for model load (#7284 )" (#7334 ) This reverts commit `583fd6b000`.	2024-05-16 20:43:45 +02:00
Radoslav Gerganov	fe34112740	rpc : get available mem for the CPU backend This can be overridden with the -m command line option ref: #7293	2024-05-16 12:04:08 +03:00
Radoslav Gerganov	0b19253ad5	rpc : add command line arg for specifying backend memory ref: #7293	2024-05-16 09:58:29 +03:00
Jared Van Bortel	7f9698470d	convert : get general.name from model dir, not its parent (#5615 ) Co-authored-by: Brian <mofosyne@gmail.com>	2024-05-16 16:15:23 +10:00
Herman Semenov	e3336679b7	grammar, json, llama: replace push on emplace if it possible (#7273 )	2024-05-16 16:14:24 +10:00
Vaibhav Srivastav	756bbb6560	doc: add references to hugging face GGUF-my-repo quantisation web tool. (#7288 ) * chore: add references to the quantisation space. * fix grammer lol. * Update README.md Co-authored-by: Julien Chaumond <julien@huggingface.co> * Update README.md Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Julien Chaumond <julien@huggingface.co> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-05-16 15:38:43 +10:00
Max Krasnyansky	ab97fe155a	ci: fix bin/Release path for windows-arm64 builds (#7317 ) Switch to Ninja Multi-Config CMake generator to resurect bin/Release path that broke artifact packaging in CI.	2024-05-16 15:36:43 +10:00
Max Krasnyansky	5cc8a89c08	Add support for properly optimized Windows ARM64 builds with LLVM and MSVC (#7191 ) * logging: add proper checks for clang to avoid errors and warnings with VA_ARGS * build: add CMake Presets and toolchian files for Windows ARM64 * matmul-int8: enable matmul-int8 with MSVC and fix Clang warnings * ci: add support for optimized Windows ARM64 builds with MSVC and LLVM * matmul-int8: fixed typos in q8_0_q8_0 matmuls Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * matmul-int8: remove unnecessary casts in q8_0_q8_0 --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-05-16 12:47:36 +10:00
Daniel Bevenius	38ecedff95	readme : remove stray double quote (#7310 ) Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2024-05-15 23:41:03 +02:00
kunnis	358f7b28cf	ggml : use dynamic thread scheduling for matrix multiplication (#6915 ) * Just reordering some structs. * Adding in the calls to mm_pause * Passing around the state * Renaming and moving a bunch of variables around. * Extracting the logic to it's own function. * Moving some variable definitions into the chunk function. * Moving some variables around * moving src1_cont inside * Moving row_size * adding the current_chunk * Reorg the code. * Formatting to match the orig patch * starting to setup the chunking variables * Starting the buildup of the loop * The yield shouldn't be necessary. * adding the looping structure based on the chunk configuration. * Add in the re-chunking code. * Making it much more likely to rechunk. * disable resizing if numa is enabled. * Updating comments with what we've learned. * Fix formatting * Couple more formatting fixes. * More style fixes. * Fix Warnings * Going with unused because there's conditional logic that needs it. * Update ggml.c * Update ggml.c ---------	2024-05-15 19:59:12 +02:00
agray3	06151b32d7	Avoid unnecessarily disabling CUDA graphs (#7302 ) As discussed in PR #6766, CUDA graphs were being disabled in the presence of long prompts. This fixes the issue by avoiding the consective update counter from incrementing unnecessarily for tokens in which cuda graphs are disabled due to batch size > 1.	2024-05-15 15:44:49 +02:00
slaren	f3f290e25e	ggml : tag ggml_tensor::backend as deprecated (#7290 )	2024-05-15 15:08:48 +02:00
AidanBeltonS	946893d257	Add missing " (#7303 )	2024-05-15 17:56:30 +05:30
dm4	0b96b615fc	embedding : free the batch after execution (#7297 )	2024-05-15 15:01:12 +03:00
Georgi Gerganov	ef6181c079	sync : ggml	2024-05-15 13:23:41 +03:00
John Balis	2ac4952938	ggml : add `ggml_upscale_ext` (ggml/814) * initial commit with CPU implementation of upscale to shape and test, cuda implementation next * experimental commit to see if dst shape is correct * test version * test * removed unnecessary params * refactor * fixed tests * ggml : metal impl + cleanup + sycl dev warnings * patched ggml_upscale cuda op to handle non-contiguous tensors, added test for non-contiguous behavior * metal : fix upsacle op to support nb00 + style --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-05-15 13:23:33 +03:00
Johannes Gäßler	cbd77f7974	server bench: fix bench not waiting for model load (#7284 )	2024-05-15 08:44:16 +02:00
Georgi Gerganov	fd3a0965b2	script : sync ggml-rpc	2024-05-14 19:14:38 +03:00
Georgi Gerganov	926059600a	metal : support FA without mask + add asserts (#7278 ) * ggml : fa without mask + add asserts ggml-ci * metal : support non-contiguous KV ggml-ci	2024-05-14 19:09:30 +03:00
Georgi Gerganov	b0390e32cf	sync : ggml ggml-ci	2024-05-14 19:08:09 +03:00
Georgi Gerganov	93c6680603	metal : tune soft_max number of threads (whisper/0)	2024-05-14 19:08:09 +03:00
Georgi Gerganov	584ab7fbfb	ggml : try fix ppc64 (whisper/0)	2024-05-14 19:08:09 +03:00
Przemysław Pawełczyk	7fdf218783	ggml : expose SSE3 and SSSE3 for MSVC when AVX is available (whisper/2128)	2024-05-14 19:08:09 +03:00
Hong Bo PENG	53475f887c	ggml : optimize for ppc64le using VSX intrinsics (ggml/784) * optimize for ppc64le using VSX intrinsics * 1. code clean up by removing comments about overflow concern. 2. fix typo in suffix of scaling. * Continue to fix typo in suffix of scaling for QK_K <> 256 --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-05-14 19:08:09 +03:00
Steve Grubb	ef87ca916e	server: free sampling contexts on exit (#7264 ) * server: free sampling contexts on exit This cleans up last leak found by the address sanitizer. * fix whitespace * fix whitespace	2024-05-14 16:11:24 +02:00
Brian	e115936501	Revert "move ndk code to a new library (#6951 )" (#7282 ) This reverts commit `efc8f767c8`.	2024-05-14 16:10:39 +03:00
Radoslav Gerganov	af81b28dbf	ggml : add RPC backend (#6829 ) * ggml : add RPC backend The RPC backend proxies all operations to a remote server which runs a regular backend (CPU, CUDA, Metal, etc). * set TCP_NODELAY * add CI workflows * Address review comments * fix warning * implement llama_max_devices() for RPC * Address review comments * Address review comments * wrap sockfd into a struct * implement get_alignment and get_max_size * add get_device_memory * fix warning * win32 support * add README * readme : trim trailing whitespace * Address review comments * win32 fix * Address review comments * fix compile warnings on macos	2024-05-14 14:27:19 +03:00
slaren	69afafd1e7	llama : disable pipeline parallelism with nkvo (#7265 )	2024-05-14 17:33:42 +10:00
Elton Kola	8536fbb140	move ndk code to a new library (#6951 )	2024-05-14 17:30:30 +10:00
Haggai Nuchi	34dbd5ac9f	Add left recursion check: quit early instead of going into an infinite loop (#7083 ) * Add left recursion check: quit early instead of going into an infinite loop * Remove custom enum, rename left recursion check and move to "grammar internal" section, add handling for edge case where a leftmost nonterminal may be empty * Remove unnecessary declaration	2024-05-14 15:25:56 +10:00
Ryuei	dc91e1430e	docs: Fix typo and update description for --embeddings flag (#7026 ) - Change '--embedding' to '--embeddings' in the README - Update the description to match the latest --help output - Added a caution about defining physical batch size	2024-05-14 15:20:47 +10:00
compilade	2ea6201d71	convert-hf : support direct Q8_0 conversion (#7234 ) * convert-hf : support q8_0 conversion * convert-hf : add missing ftype This was messing with the checksums otherwise. * convert-hf : add missing ftype to Baichuan and Xverse I didn't notice these on my first pass.	2024-05-13 14:10:51 -04:00
Georgi Gerganov	b60d93f7f7	llama : less KV padding when FA is off (#7257 ) ggml-ci	2024-05-13 17:15:15 +03:00
k.h.lai	e0baf1aca7	llava-cli: fix base64 prompt (#7248 )	2024-05-14 00:02:36 +10:00
Johannes Gäßler	2d2147923e	perplexity: add BF16 vs. FP16 results (#7150 )	2024-05-13 13:03:27 +02:00
Neo Zhang	b97711d6c1	[SYCL] rm wait() (#7233 )	2024-05-13 18:11:26 +08:00
Joan Fontanals	bf009f1d45	llama : rename jina tokenizers to v2 (#7249 ) * refactor: rename jina tokenizers to v2 * refactor: keep refactoring non-breaking	2024-05-13 11:35:14 +03:00
Brian	97684629fa	convert.py: Outfile default name change and additional metadata support (#4858 ) * convert.py: Outfile default name change and additional metadata support * convert.py: don't stringify Metadata load method output * convert.py: typo fix * convert.py: fix metadata format to sync with LLM_KV_NAMES in llama.cpp	2024-05-13 12:56:47 +10:00
Benjamin Findley	9b9b4eb946	change default temperature of OAI compat API from 0 to 1 (#7226 ) * change default temperature of OAI compat API from 0 to 1 * make tests explicitly send temperature to OAI API	2024-05-13 12:40:08 +10:00
Neo Zhang	1fa2d4319b	[SYCL] Add oneapi runtime dll files to win release package (#7241 ) * add oneapi running time dlls to release package * fix path * fix path * fix path * fix path * fix path --------- Co-authored-by: Zhang <jianyu.zhang@intel.com>	2024-05-13 08:04:29 +08:00
Neo Zhang	b079bf29ed	[SYCL] update CI with oneapi 2024.1 (#7235 ) Co-authored-by: Zhang <jianyu.zhang@intel.com>	2024-05-13 08:02:55 +08:00

1 2 3 4 5 ...

2912 Commits