ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-02-01 20:19:52 +00:00

Author	SHA1	Message	Date
staviq	380d2fca81	server : support for saving templates in browser LocalStorage (#2486 ) * support for templates in browser LocalStorage * sync accepted #2409 fix from upstream * convert autosave invocation to useEffect * Apply suggestions from code review Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com> * Regen index.html.cpp, suggested from code review --------- Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com>	2023-08-18 07:34:01 +08:00
Kerfuffle	aad3ce7fee	Add --cfg-negative-prompt-file option for examples (#2591 ) Add --cfg-negative-prompt-file option for examples	2023-08-17 07:29:44 -06:00
Jhen-Jie Hong	87be3696f6	server : add missing /json-schema-to-grammar.mjs (#2616 ) fixes #2611	2023-08-15 06:14:14 +08:00
Cheng Shao	bb04b1c694	server : add --numa support (#2524 )	2023-08-14 16:36:42 +03:00
Jhen-Jie Hong	16aaeef5b3	server : fix default grammar by use empty string in the UI (#2604 )	2023-08-14 16:20:17 +08:00
Jhen-Jie Hong	4fce312113	server : implement json-schema-to-grammar.mjs & add grammar param in the UI (#2588 ) * server : implement json-schema-to-grammar.mjs by follow python impl * server : add grammar support in chat.mjs * server : implement grammer param in the UI * server : generate .hpp * server : remove trailing whitespaces * server : generate .hpp * server : fix sort of prop pairs * server : optimize regex & iteration	2023-08-14 15:16:54 +08:00
byte-6174	bb9ebb4394	Adding support for llama2.c models (#2559 )	2023-08-12 01:17:25 +02:00
Equim	ae2377983b	server: fixed wrong variable name in timing json (#2579 ) * server: fixed wrong variable name in timing json * remove redunct entry	2023-08-12 00:35:14 +02:00
DannyDaemonic	17defbc44e	Handle `ENABLE_VIRTUAL_TERMINAL_PROCESSING` more gracefully on earlier versions of Windows.	2023-08-10 13:11:36 -07:00
Christian Demsar	c366b01720	Add --n-predict -2 for stopping generation on full context (#2565 )	2023-08-10 16:28:27 +02:00
Martin Krasser	17d5b4bfbe	Fix grammar-based sampling issue in server (#2566 )	2023-08-10 13:16:38 +03:00
Martin Krasser	220649a37f	Allow passing grammar to completion endpoint (#2532 ) * Allow passing grammar to completion endpoint	2023-08-08 16:29:19 +03:00
chaihahaha	932dac4a50	llm.vim : multiline autocompletion, get rid of "^@" (#2543 )	2023-08-08 15:07:02 +03:00
Georgi Gerganov	74f72ea675	vim : bring back simple llm.vim example	2023-08-08 15:06:18 +03:00
AustinMroz	d517b2f07c	vim : streaming and more (#2495 ) * Update Vim plugin * Remove getbufoneline usage, Add input bind example. getbufoneline() appears to be a recently added function and has been replaced with getbufline for compatibility. An additional example that explains how to add a keybind that works in insert mode was added.	2023-08-08 14:44:48 +03:00
klosax	fb17afa861	Add --rope-scale parameter (#2544 ) * common.cpp : Add --rope-scale parameter * README.md : Add info about using linear rope scaling	2023-08-07 19:07:19 +02:00
DannyDaemonic	28fe11b65e	console : fix issue related to Windows 11 PowerShell console mode persistence (#2521 )	2023-08-06 09:49:34 +03:00
Jonas Wunderlich	e45dbc8951	fix firefox autoscroll (#2519 )	2023-08-04 22:16:11 +02:00
Cebtenzzre	6ae47d1968	server: regenerate completion.js.hpp (#2515 )	2023-08-04 21:00:57 +02:00
DannyDaemonic	b1cabde3b3	Add --simple-io option for subprocesses and break out console.h and cpp (#1558 )	2023-08-04 08:20:12 -07:00
Stephen Nichols	d1d9a25cce	Fixing race condition in server and partial stream handling in frontend. (#2391 ) * Fixing race condition in server.cpp and partial stream handling in completion.js * Reverting assert edits. * Adding newline to eof	2023-08-04 13:37:24 +02:00
Borislav Stanimirov	04e64ce94e	build : fix several cast and printf warnings (#2499 )	2023-08-04 13:07:21 +03:00
Evan Jones	79f99c6b4d	examples : generate JSON according to schema (#1887 ) * examples : add JSON schema grammars * complete JSON grammar * ensure primitive types can be used as root of schema * support integer type and adjust usage text	2023-08-02 22:05:44 -04:00
Eve	31e5af188f	tests : Fix compilation warnings (Linux/GCC) (#2451 ) * fix hellaswag print format, cast away warning in test-double-float * c++11 cannot use designated initializers * add static to test-grad0.c internal functions * use memcpy in test-double-float.c * port c tests to c++ * use initializer list for ggml_init_params	2023-08-02 11:06:19 +03:00
Bono Lv	c8d9f77356	fix a typo in examples/server/README.md (#2478 )	2023-08-01 14:54:28 +02:00
ebraminio	7386a7dd10	server : Support dark mode (#2414 ) * server : Support dark mode So it respects user system light / dark settings. * Update index.html.hpp by running ./deps.sh	2023-08-01 10:56:23 +02:00
Johannes Gäßler	83772f01d7	CUDA: mmq CLI option, fixed mmq build issues (#2453 )	2023-07-31 15:44:35 +02:00
klosax	fdcc7db2a5	perplexity : add Hellaswag calculation (#2389 ) * common.h : add hellaswag / remove perplexity-lines * common.cpp : add hellaswag / remove perplexity-lines * perplexity.cpp : add hellswag scores / remove perplexity-lines * perplexity.cpp : clean up * common.h : change default param value * common.cpp : Change default param * perplexity.cpp : alter wording * common.h : alter wording * common.cpp : alter wording	2023-07-28 21:25:36 +03:00
Georgi Gerganov	ca55f98a26	examples : fix whitespace	2023-07-28 21:05:08 +03:00
nhamanasu	01a14d2c58	examples : server chat mode with llama2 (#2400 ) * add: server chat mode with llama2 * fix: remove the unnecessary last \n	2023-07-28 21:02:10 +03:00
Weird Constructor	fb09f01651	readme : fix the description of the Tail free sampling (TFS) method (#2431 )	2023-07-28 11:44:43 +03:00
Rand Xie	0c7137b4c3	llama : use n_embd_gqa instead of n_embd to handle llama-2 70B (#2433 )	2023-07-28 11:42:53 +03:00
Kawrakow	abd493a08a	Add LLAMA_DEFAULT_RMS_EPS so we can change the default (#2384 ) Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>	2023-07-25 18:35:53 +03:00
Xiao-Yong Jin	fdfc76e541	main : add `--in-prefix-bos` to prefix BOS to user inputs; keep EOS (#2304 ) * add `--in-prefix-bos` to prefix BOS to user inputs; keep EOS The BOS precedes the string specified by `--in-prefix`. Model generated EOS is now kept in the context. It provides a way to strictly following the prompt format used in Llama-2-chat. The EOS handling also benefits some existing finetunes that uses EOS to mark the end of turn. * examples/common: move input_prefix_bos to other bools	2023-07-25 15:19:11 +03:00
slaren	8b710f6c6c	server: add rms_norm_eps parameter (#2380 )	2023-07-25 12:36:17 +03:00
Henri Vasserman	85913309c2	[Server] Escape HTML in webchat (#2368 ) * escape HTML in webchat * add amp	2023-07-25 10:27:34 +03:00
slaren	f03ca6ee8a	make rms_norm_eps a parameter (#2374 ) * make rms_norm_eps a parameter * add rms_norm_eps to command line * fix baby llama, test-grad0 * use scientific notation for eps param in the help ggml-ci	2023-07-24 17:57:12 +02:00
Aarni Koskela	6d91bf5e54	Chat UI extras (#2366 ) * makefile: correct deps for server * server: tighten settings layout a little * server: expose all currently configured generation params in UI * server: expose remaining generation params, for the adventurous * server: embetter mirostat fields	2023-07-24 17:54:22 +03:00
Evan Jones	9e7ca68b90	llama : add grammar-based sampling (#1773 ) * llama, main : constrain sampling to grammar * allow loading grammar from file * fix whitespace errors * handle & print parser errors * add comments to grammar syntax and allow newlines where unambiguous * add missing include * support alternates in root rule * fix bugs with empty token and EOS * adjust JSON grammar * remove swp file * rewrite ternary expressions Co-authored-by: Henri Vasserman <henv@hot.ee> * use struct for grammar elements and add Unicode support * add unicode escapes * add inverse char ranges * only sample full tokens (no peeking or truncation) * llama : minor style changes blindly applied in online editor - hopefully I didn't break something * update help text * add warning message if EOS is disabled --------- Co-authored-by: Henri Vasserman <henv@hot.ee> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-23 23:58:10 -04:00
IgnacioFDM	c4a2ed79be	Add gqa parameter support to the server (#2351 ) * Add gqa parameter support to the server * Change help from stderr to stdout	2023-07-23 23:31:17 +03:00
wzy	f65f81e723	common : n_threads == -1 uses std:🧵:hardware_concurrency() (#2347 ) * Fix #2345, fix incorrect n_threads * Update examples/common.cpp --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-23 16:33:02 +03:00
Georgi Gerganov	3fe463b26d	llama : grouped-query attention + LLaMAv2 70B support (#2276 ) * CUDA: GQA implementation * llama : support for GQA and LLaMAv2 70B ggml-ci * py : fix hparams parsing (if-else blocks) ggml-ci * py : oh boy .. ggml-ci * help : fix gqa value for 70B ggml-ci --------- Co-authored-by: JohannesGaessler <johannesg@5d6.de>	2023-07-23 15:09:47 +03:00
maddes8cht	02086139f9	llama : print help to stdout (#2338 )	2023-07-23 14:59:48 +03:00
AustinMroz	574ccaa8b8	examples : simplify vim plugin (#2327 ) Uses builtin json_encode and json_decode functions to simplify escaping Removes the need for temp files	2023-07-23 14:16:48 +03:00
Georgi Gerganov	926b090575	llama : optimize memory buffers (#2325 )	2023-07-22 21:17:57 +03:00
klosax	0c4de4eb95	Perplexity: Compute scores correlated to HellaSwag (#2312 ) * Add parameter --perplexity-lines to perplexity.cpp	2023-07-22 14:21:24 +02:00
whoreson	6f5d82363e	examples : basic VIM plugin VIM plugin for server exe	2023-07-22 13:34:51 +03:00
Richard Roberson	d20efa0a13	examples : add easy python script to create quantized (k-bit support) GGML models from local HF Transformer models (#2311 ) * Resync my fork with new llama.cpp commits * examples : rename to use dash instead of underscore --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-07-21 22:01:10 +03:00
Ikko Eltociear Ashimine	0f8aca4e96	examples : fix typo in minigpt4.py (#2298 ) promt -> prompt	2023-07-21 14:53:07 +03:00
Georgi Gerganov	b3b8017933	ggml : fix rope args order + assert (#2054 )	2023-07-21 14:51:34 +03:00

1 2 3 4 5

234 Commits