ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-02-03 04:57:32 +00:00

Author	SHA1	Message	Date
Kawrakow	ab18bc5e87	k_quants tuning for Falcon-7b (#2816 ) * Make ggml-cuda.cu build with QK_K = 64 Using LLAMA_CUDA_FORCE_DMMV = ON and -nommq it runs and produces a meaningful result. * k_quants tuning for Falcon-7b --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2023-08-27 15:19:59 +03:00
Georgi Gerganov	5a7aaa5f74	readme : update hot topics	2023-08-27 14:44:35 +03:00
Georgi Gerganov	11f0bd7499	gguf : add 64-bit support (GGUF v2) (#2821 ) * gguf : bump version to 2 * gguf : add support for 64-bit (no backwards comp yet) * gguf : v1 backwards comp * gguf.py : bump GGUF version * gguf.py : uint64_t on all lengths, sizes and counts, enums still uint32_t * gguf.py : string lengths uint32_t * gguf : update all counts to 64-bit * gguf.py : string len uint64_t and n_dims uint32_t * gguf : fix typo * llama.cpp : print gguf version --------- Co-authored-by: klosax <131523366+klosax@users.noreply.github.com>	2023-08-27 14:19:54 +03:00
Georgi Gerganov	926dbcbaab	llama : more tokenizer fixes (#2810 ) * tests : write a Python tokenizer test (wip) * llama : prefix input text for tokenization with whitespace * llama : distinguish pieces from decoded text + fix detokenization * common : add comments * examples : no longer manually add leading space when tokenizing * tests : use Python to generate tokenizer tests for C++ * tests : add option to tokenize text files ggml-ci * tests : add test-tokenizer-1.py * llama.cpp : fix LF token * hellaswag : move the concat space for clarity * tests : add falcon tests (py + cpp, currently do not pass Unicode) ggml-ci * common : temporary separate llama_detokenize calls for SPM and BPE --------- Co-authored-by: klosax <131523366+klosax@users.noreply.github.com>	2023-08-27 14:19:19 +03:00
Przemysław Pawełczyk	0c3eafbe0e	ggml : detect SSSE3 (#2825 ) * ggml : add ggml_cpu_has_ssse3 * llama : show SSSE3 in system info	2023-08-27 11:10:25 +03:00
slaren	699ba791e5	ci : add LoRA test to CI (#2650 ) * ci : add lora test ggml-ci * move lora summary to the top, add lora logs ggml-ci * ci : decrease CPU ppl runs to 2 to avoide 20 min timeout ggml-ci * add 7b lora test use 1 thread for CUDA generation tests ggml-ci * add test with q8_0 (cpu only) ggml-ci --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-08-27 10:03:27 +03:00
Bruce MacDonald	f2256c49e0	server : add `/detokenize` endpoint (#2802 ) * Add a /detokenize endpoint to the example server * remove trailing white-space	2023-08-27 07:11:45 +08:00
Kerfuffle	a84214bd70	convert.py : advanced option (#2753 ) * Allow convert.py to convert to q8_0 Fix issue with bounded_parallel_map and greedy consuming iterator Display elapsed time during conversion * Add --concurrency option Minor improvements to help text Clean up bounded_parallel_map function a bit * Massive speed improvement thanks to Cebtenzzre * Refactor types	2023-08-26 23:13:36 +03:00
Tim Miller	e604afedc8	llama : use Unicode Escape Sequence to replace encoded characters (#2814 ) The use of special characters within source files can break compiling on some computers with different region and language settings. Using Unicode escape sequences should allow for the code to be compiled on all setups without needing to change your computers settings or switch regions.	2023-08-26 21:27:07 +03:00
Tungsten842	10f684465a	flake.nix : add rocm support and cleanup (#2808 )	2023-08-26 21:19:44 +03:00
Cebtenzzre	03eea5f437	llama : move #includes out of _GNU_SOURCE conditional (#2817 )	2023-08-26 21:17:51 +03:00
Dr. Tom Murphy VII Ph.D	92a41f322b	main : fix bug (penalize_nl=false doesn't work) + suppress warning on mingw (#1528 ) * Fix bug in main.cpp where penalize_nl=false has no effect. It modifies the underlying logits array, but at this point we are already working on the candidates copy. * Suppress redefinition warning for NOMINMAX on mingw. In my installation, this macro is already defined by /usr/lib/gcc/x86_64-w64-mingw32/11/include/c++/x86_64-w64-mingw32/bits/os_defines.h:45. * main : fix indentation * main : pass ctx to llama_token_nl() --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-08-26 21:12:56 +03:00
Cebtenzzre	ea11e400d7	llama : use std::abs in llama_sample_tail_free (#2800 ) Plain 'abs' casts the input to int.	2023-08-26 19:53:52 +03:00
Georgi Gerganov	d821a8f3f7	k-quants : remove unnecessary tensor shape restrictions (#2811 )	2023-08-26 17:37:35 +03:00
Kawrakow	9b782d829a	Better perplexity for 2- and 3-bit quantization for LLaMA-v2-70B (#2807 ) * Better perplexity for 2- and 3-bit quantization for the 70B model * PR comment --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2023-08-26 17:27:49 +03:00
Kawrakow	756deeaf3c	Fix HellaSwag (#2805 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2023-08-26 16:48:53 +03:00
Volodymyr Vitvitskyi	3f42f52f01	flake : build llama.cpp on Intel with nix (#2795 ) Problem ------- `nix build` fails with missing `Accelerate.h`. Changes ------- - Fix build of the llama.cpp with nix for Intel: add the same SDK frameworks as for ARM - Add `quantize` app to the output of nix flake - Extend nix devShell with llama-python so we can use convertScript Testing ------- Testing the steps with nix: 1. `nix build` Get the model and then 2. `nix develop` and then `python convert.py models/llama-2-7b.ggmlv3.q4_0.bin` 3. `nix run llama.cpp#quantize -- open_llama_7b/ggml-model-f16.gguf ./models/ggml-model-q4_0.bin 2` 4. `nix run llama.cpp#llama -- -m models/ggml-model-q4_0.bin -p "What is nix?" -n 400 --temp 0.8 -e -t 8` Co-authored-by: Volodymyr Vitvitskyi <volodymyrvitvitskyi@SamsungPro.local>	2023-08-26 16:25:39 +03:00
Nigel Bosch	ac6575162d	Handle null rope scaling value (#2793 )	2023-08-26 14:11:17 +02:00
klosax	976b621020	Fix spm whitespaces (#2806 ) * llama.cpp : fix spm whitespace escaping + clean up * main.cpp : spm - add whitespace in front of prompt * test-tokenizer-0.cpp : spm - add whitespace in front of prompt	2023-08-26 13:45:53 +02:00
lon	7ede4319fc	examples : skip unnecessary external lib in server README.md how-to (#2804 )	2023-08-26 16:07:43 +08:00
Marcus Dunn	f19ed06ed0	llama : fix struct decl (#2790 )	2023-08-25 19:17:15 +03:00
Kawrakow	198140c2aa	Faster perplexity computation (#2786 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2023-08-25 19:05:02 +03:00
Matt Pulver	3e0b38e027	llama : add llama_beam_search() (#2267 ) * Add llama_beam_search(). * Add '// Beam search' heading to llama.{h,cpp} after llama_grammar_accept_token(). * Add space around * pointers and & references. * Add spaces around comparison and assignment operators. * Prefer west const. * Use llama_ prefix for structs in global namespace. * Delete obsolete comment from an earlier revision. * Change eos to eob in llama_beam and llama_beam_view structs.	2023-08-25 18:18:48 +03:00
Nigel Bosch	8ef2a0c9d3	convert.py : Get rope scale from HuggingFace models (#2772 ) * Get rope scale from HF models * Save rope scale only for linear scaling * Rewrite for clarity	2023-08-25 16:41:52 +02:00
slaren	89e4a4461e	llama-bench : add model sizes (#2771 ) * llama-bench : add model sizes * more compact markdown output * back to GiB * adjust column sizes	2023-08-25 15:16:19 +02:00
slaren	e20b657ffb	convert.py : export rope freq_base when converting CodeLlama from an HF model (#2773 )	2023-08-25 14:08:53 +02:00
Jhen-Jie Hong	b06380dcc0	server : display token probabilities in the UI (#2489 ) * server : add n_probs param in chat UI * server : keep message data array & show in probabilites component * server : add simple popover component * server : fix completion_probabilities undefined if not set n_probs * server : implement Probabilites * server : handle bytes * server : make n_probs max to 10 for easy scroll * server : adjust for dark/light mode * server : Fix regenerated prompt * server : update index.html.hpp * server : convert prob to percentage + show original value as div title * server : fix Probabilites not used if included empty str * server : skip byte pair in display probabilites * server : remove array check of completion_probabilities in messages * skip empty array or byte pair (> 1) in Probabilites * generate index.html.hpp * fix incorrect prob convert if the str is already a known token * use final response to show probabilities on stop * revert unnecessary change * correct probabilites usage * remove unused function * always send partial response for get correct probs of last to_send * fix typo * fix content of format_final_response * refactor probs render & make pColor transparent if not found * send empty string when got stop_pos in partial * avoid unnecessary empty data event & send rest of partial tokens on stop * use <br /> for new line * skip -1 tok in loop to avoid send '' on end * trim last new lines on stop * revert unnecessary change	2023-08-25 18:32:45 +08:00
Georgi Gerganov	6224f81799	ci : pip install gguf in editable mode (#2782 ) ggml-ci	2023-08-25 13:03:25 +03:00
M. Yusuf Sarıgöz	08a1012230	gguf : export objects to user code (#2780 ) * gguf export more objects to user code * gguf export all objects to user code for now * gguf : bump version	2023-08-25 12:43:41 +03:00
Henri Vasserman	984b7495ed	ROCm Port (#1087 ) * use hipblas based on cublas * Update Makefile for the Cuda kernels * Expand arch list and make it overrideable * Fix multi GPU on multiple amd architectures with rocblas_initialize() (#5) * add hipBLAS to README * new build arg LLAMA_CUDA_MMQ_Y * fix half2 decomposition * Add intrinsics polyfills for AMD * AMD assembly optimized __dp4a * Allow overriding CC_TURING * use "ROCm" instead of "CUDA" * ignore all build dirs * Add Dockerfiles * fix llama-bench * fix -nommq help for non CUDA/HIP --------- Co-authored-by: YellowRoseCx <80486540+YellowRoseCx@users.noreply.github.com> Co-authored-by: ardfork <134447697+ardfork@users.noreply.github.com> Co-authored-by: funnbot <22226942+funnbot@users.noreply.github.com> Co-authored-by: Engininja2 <139037756+Engininja2@users.noreply.github.com> Co-authored-by: Kerfuffle <44031344+KerfuffleV2@users.noreply.github.com> Co-authored-by: jammm <2500920+jammm@users.noreply.github.com> Co-authored-by: jdecourval <7315817+jdecourval@users.noreply.github.com>	2023-08-25 12:09:42 +03:00
Georgi Gerganov	40c8c6dd6f	cuda : add RoPE kernel for mode == 2 (NeoX) (#2760 ) * cuda : add RoPE kernel for mode == 2 (NeoX) * falcon : do not offload the embeddings layer	2023-08-25 11:55:59 +03:00
M. Yusuf Sarıgöz	a67ec14fe9	gguf : make gguf pip-installable * gitignore : add dist and rm pyproject.toml * gguf: prepare as Pip package * gguf: prepare as Pip package * gguf : fix line endings * requirements : add gguf * gguf : update readme with build notes * gguf : update readme with build notes * gguf : add notes for tests	2023-08-25 09:26:05 +03:00
Shouzheng Liu	800ef93db4	ggml-alloc : enlarge size of parse_seq (#2776 ) Since we also store barriers in this array, we need to double its size.	2023-08-25 08:58:00 +03:00
Marcus Dunn	1e8200c3d0	Added `enum` to `llama_token_get_type` return type (#2774 )	2023-08-24 23:49:30 +02:00
slaren	506fd81d05	convert.py : try to determine n_ctx automatically for CodeLlama (#2770 )	2023-08-24 21:10:39 +02:00
slaren	9818be3377	gguf : add rope_freq_base parameter for CodeLlama (#2769 )	2023-08-24 21:04:05 +03:00
Georgi Gerganov	9042737101	falcon : write file type	2023-08-24 19:58:30 +03:00
Shouzheng Liu	dc51c17e4c	metal : bug-fix when enable ggml-alloc (#2757 ) * metal: better memory alloc w/ concurrency dispatch The ggml-alloc should only free tensors at memory barriers. * ggml-alloc: avoid return silently In certain cases, the allocate_node() function may silently return without performing any memory allocation.	2023-08-24 19:27:25 +03:00
Georgi Gerganov	8b08abe24f	convert : auto-determine model name based on dir + scripts update	2023-08-24 19:26:47 +03:00
Kerfuffle	2a2645fd76	Fix for main example getting stuck when -n -2 and --interactive (#2767 ) * Fix for main example getting stuck when -n -2 and --interactive * Add a comment so future generations may suffer less.	2023-08-24 10:11:13 -06:00
slaren	3b743a5340	fix convert.py for codellama, add llama 34B to the list of recognized models (#2768 )	2023-08-24 17:44:11 +02:00
DannyDaemonic	a74a205f64	Tag release with build number (#2732 ) * Modified build.yml to use build number for release * Add the short hash back into the tag * Prefix the build number with b	2023-08-24 15:58:02 +02:00
Georgi Gerganov	25399c1197	metal : add Q8_0 support (#2763 ) * metal : add dequantize_q8_0 kernel * metal : add mul_mat_q8_0_f32 kernel * metal : add Q8_0 mul_mm kernel	2023-08-24 16:19:57 +03:00
Georgi Gerganov	96e9fad81f	llama : escape all U+2581 in a string (#2750 )	2023-08-24 12:26:01 +03:00
Evan Jones	f4102e260a	llama : fix grammar sometimes generating null char (#2756 )	2023-08-24 07:07:13 +03:00
Georgi Gerganov	fc84c48240	readme : fix link	2023-08-23 23:44:19 +03:00
Georgi Gerganov	1fac3b2c0b	minor : fix trailing whitespace	2023-08-23 23:43:00 +03:00
Georgi Gerganov	eb5bf4480c	readme : update hot topics	2023-08-23 23:41:16 +03:00
Georgi Gerganov	5faba0e8a3	llm : add Falcon support (#2717 ) * llama : refactor GGUF constants into static maps * llama : check if model architecture is known * llama : refactor llama_model_load_internal() * gguf : add KV constant maps * llm : read arch-specific KVs * convert : add dummy scores + types * falcon : load tensor data (CPU only) * llama : fix loading progress bar * llama : add arch member to llama_model * falcon : CPU inference working * falcon : support non-40B models * falcon : minor * llama : minor updates ggml-ci * convert-falcon-hf-to-gguf.py : fix special token mapping * llama.cpp : llama default UNK token = id 0 * llama.cpp : fix bpe tokenizer * llama.cpp : fix the fix of bpe tokenizer * ggml : pass eps to ggml_norm * metal : implement RoPE (mode = 2) + avoid ggml_repeat * ggml : ggml_repeat always creates new tensor * falcon : copy-paste self-attention from LLaMA * metal : print extra compute pipeline info * falcon : minor changes (still chasing the Metal problem) * llama.cpp : fix linefeed token * metal : fix GELU kernel numerical stability by using precise::tanh * metal : temporary workaround for the concurrency optimization bug * falcon : add CUDA offloading (#2739) * llama : better model naming and size reporting * llama : prep new tokenizer support * llama : advanced BPE tokenizer based on ggllm.cpp imlpementation * llama : remove oboslete comment ggml-ci * common : remove obsolete BPE API + disable test-tokenizer-1 * llama : revert BPE special-case in llama_byte_to_token() * cuda : add TODOs for RoPE NeoX implementation * llama : default special tokens based on vocab type * perplexity : add log for start of tokenization --------- Co-authored-by: klosax <131523366+klosax@users.noreply.github.com> Co-authored-by: slaren <slarengh@gmail.com>	2023-08-23 23:08:04 +03:00
Georgi Gerganov	4e35149e03	minor : fix trailing whitespace	2023-08-23 22:37:39 +03:00

... 30 31 32 33 34 ...

2639 Commits