merge from upstream #20

l3utterfly · 2024-05-22T06:36:33Z

No description provided.

…#7284)" (ggerganov#7334) This reverts commit 583fd6b.

This change upstreams llamafile's vectorized expf() functions. This lets us compute softmax and silu more accurately than the short[65536] lookup table that GGML previously used to make this operation go faster. We can support aarch64 and sse2+ with the worst case rounding error of 2ulp. It makes make -j8 tests && ./tests/test-backend-ops -o SOFT_MAX -b CPU perf go 1.5x faster for SSE2+FMA, 1.9x faster for AVX2+FMA and 2.1x on AVX512

ref: ggerganov#7292

* llama : use n_embd_head_v instead of n_embd_head_k when reshaping kqv * llama : use n_embd_v_gqa and n_embd_head_v instead of n_embd_k_gqa and n_embd_head_k when making a view of cached value vectors. --------- Co-authored-by: Stanisław Szymczyk <[email protected]>

* convert-hf-to-gguf-update: automate updating * convert-hf-to-gguf-update: improve download * share requests session for performance * create directories only when needed, don't skip downloads when empty directory encountered * be more graceful about errors

…robust (ggerganov#7279) * run-single-test.sh: added a single test function script and fix debug-test.sh to be more robust * debug-test.sh: combined execute and gdb test mode via -g flag * debug-test.sh: refactor * debug-test: refactor for clarity * debug-test.sh: comment style changes * debug-test.sh: fix gdb

ref: ggerganov#7293

@jin-eld

Supercedes ggerganov#4024 and ggerganov#4813. CMake's native HIP support has become the recommended way to add HIP code into a project (see [here](https://rocm.docs.amd.com/en/docs-6.0.0/conceptual/cmake-packages.html#using-hip-in-cmake)). This PR makes the following changes: 1. The environment variable `HIPCXX` or CMake option `CMAKE_HIP_COMPILER` should be used to specify the HIP compiler. Notably this shouldn't be `hipcc`, but ROCm's clang, which usually resides in `$ROCM_PATH/llvm/bin/clang`. Previously this was control by `CMAKE_C_COMPILER` and `CMAKE_CXX_COMPILER`. Note that since native CMake HIP support is not yet available on Windows, on Windows we fall back to the old behavior. 2. CMake option `CMAKE_HIP_ARCHITECTURES` is used to control the GPU architectures to build for. Previously this was controled by `GPU_TARGETS`. 3. Updated the Nix recipe to account for these new changes. 4. The GPU targets to build against in the Nix recipe is now consistent with the supported GPU targets in nixpkgs. 5. Added CI checks for HIP on both Linux and Windows. On Linux, we test both the new and old behavior. The most important part about this PR is the separation of the HIP compiler and the C/C++ compiler. This allows users to choose a different C/C++ compiler if desired, compared to the current situation where when building for ROCm support, everything must be compiled with ROCm's clang. ~~Makefile is unchanged. Please let me know if we want to be consistent on variables' naming because Makefile still uses `GPU_TARGETS` to control architectures to build for, but I feel like setting `CMAKE_HIP_ARCHITECTURES` is a bit awkward when you're calling `make`.~~ Makefile used `GPU_TARGETS` but the README says to use `AMDGPU_TARGETS`. For consistency with CMake, all usage of `GPU_TARGETS` in Makefile has been updated to `AMDGPU_TARGETS`. Thanks to the suggestion of @jin-eld, to maintain backwards compatibility (and not break too many downstream users' builds), if `CMAKE_CXX_COMPILER` ends with `hipcc`, then we still compile using the original behavior and emit a warning that recommends switching to the new HIP support. Similarly, if `AMDGPU_TARGETS` is set but `CMAKE_HIP_ARCHITECTURES` is not, then we forward `AMDGPU_TARGETS` to `CMAKE_HIP_ARCHITECTURES` to ease the transition to the new HIP support. Signed-off-by: Gavin Zhao <[email protected]>

* Replace CODEPOINT_TYPE_* with codepoint_flags * Update and bugfix brute force random test * Deterministic brute force random test * Unicode normalization NFD * Get rid of BOM

…ero (ggerganov#7313)

* convert : fix set_vocab_sentencepiece * Update convert-hf-to-gguf.py

* github-actions-labeler: initial commit [no ci] * github actions: remove priority auto labeling [no ci]

…#7237) * Update and fix Vulkan softmax implementation * Update and fix Vulkan argsort implementation

…#7348) Fix floating point error with ndot printing, allow end stats on lower task numbers if multiple-choice tasks.

…nov#7324) Tie the weights for ARCH_STARCODER to support the larger Granite code models. Partially addresses ggerganov/issues/7116 There still remains to be a few things to fix. Currently requires `--override-kv tokenizer.ggml.add_bos_token=bool:false`

* android : use "ci-android" branch for CI * ggml : disable SIMD exp and silu for 32-bit ARM ggml-ci * android : do not fetch, use add_subdirectory instead * cmake : provide binary dir

* Revert "ci : temporary disable sanitizer builds (ggerganov#6128)" This reverts commit 4f6d133. * ci : trigger

* logging: output capture in cuda module * fix compile error * fix: vsnprintf terminates with 0, string use not correct * post review * Update llama.cpp Co-authored-by: slaren <[email protected]> * Update llama.cpp Co-authored-by: slaren <[email protected]> --------- Co-authored-by: slaren <[email protected]>

…#7363) https://github.com/actions/labeler#using-configuration-path-input-together-with-the-actionscheckout-action Recommends the use of checkout action to use the correct repo context when applying settings for PR labels e.g. steps: - uses: actions/checkout@v4 # Uploads repository content to the runner with: repository: "owner/repositoryName" # The one of the available inputs, visit https://github.com/actions/checkout#readme to find more - uses: actions/labeler@v5 with: configuration-path: 'path/to/the/uploaded/configuration/file'

* Add StableLM pre-tokenizer * Fix space * Fix trailing whitespace

* Fix empty Vulkan host buffers Add fp32 fp16 matmul shader Fix matmul shader alignment * Remove deprecated tensor->backend uses * Fix Vulkan validation errors on embedding models with no offloaded layers * Fix Vulkan llava segfault when not offloading layers

…ision for enabling AVX512_BF16 (ggerganov#7258)

* server : don't pass temperature as string * server : increase timeout * tests : fix the fix 0.8f -> 0.8 ggml-ci * tests : set explicit temperature

* add loongarch lsx and lasx optimize code * Add loongarch compilation support to makefile * revert stb_image.h * opt bytes_from_nibbles_32 and sum_i16_pairs_float * fix undeclared * format code * update * update 2 --------- Co-authored-by: Jinyang He <[email protected]>

…#7272)

* Update SYCL upscale operation * Formatting * Remove messages

* server : fix temperature * server : disable tests relying on parallel determinism * ci : change server Debug -> RelWithDebInfo

* rpc : track allocated buffers ref: ggerganov#7407 * rpc : pack rpc_tensor tightly

* llama : remove Persimmon * requirements : remove

* Update brute force test: special tokens * Fix added tokens - Try to read 'added_tokens.json'. - Try to read 'tokenizer_config.json'. - Try to read 'tokenizer.json'. * Fix special tokens rtrim Co-authored-by: Georgi Gerganov <[email protected]> * server : fix test regexes

* Update brute force test: add_special * Update brute force test: default values for add_bos_token and add_eos_token * Enable rtrim when pre-inserting BOS Co-authored-by: Georgi Gerganov <[email protected]> * Revert "server : fix test regexes"

* examples: cache hf model when --model not provided * examples: cache hf model when --model not provided * examples: cache hf model when --model not provided * examples: cache hf model when --model not provided * examples: cache hf model when --model not provided

ggml-ci

* add phi3 128k support in convert-hf-to-gguf * add phi3 128k support in cuda * address build warnings on llama.cpp * adjust index value in cuda long rope freq factors * add long rope support in ggml cpu backend * make freq factors only depend on ctx size * remove unused rope scaling type 'su' frin gguf converter * fix flint warnings on convert-hf-to-gguf.py * set to the short freq factor when context size is small than trained context size * add one line of comments * metal : support rope freq_factors * ggml : update ggml_rope_ext API to support freq. factors * backends : add dev messages to support rope freq. factors * minor : style * tests : update to use new rope API * backends : fix pragma semicolons * minor : cleanup * llama : move rope factors from KV header to tensors * llama : remove tmp assert * cuda : fix compile warning * convert : read/write n_head_kv * llama : fix uninitialized tensors --------- Co-authored-by: Georgi Gerganov <[email protected]>

phymbert and others added 30 commits May 16, 2024 20:43

Revert "server bench: fix bench not waiting for model load (ggerganov…

24ecb58

…#7284)" (ggerganov#7334) This reverts commit 583fd6b.

[Server] Added --verbose option to README [no ci] (ggerganov#7335)

9c4fdcb

server : add support for the RPC backend (ggerganov#7305)

ee94172

ref: ggerganov#7292

convert : fix Qwen/Qwen-7b conversion (ggerganov#7308)

e18bc6a

ggml-quants, llama : removed excess checks (ggerganov#7274)

359cbe3

tokenization: add warning for double BOS (ggerganov#7332)

29c60d8

rpc : set SO_REUSEADDR for the server socket (ggerganov#7320)

f4bd8b3

ref: ggerganov#7293

CUDA: faster large batch FA without tensor cores (ggerganov#7314)

0fc1e82

Unicode codepoint flags for custom regexs (ggerganov#7245)

b43272a

* Replace CODEPOINT_TYPE_* with codepoint_flags * Update and bugfix brute force random test * Deterministic brute force random test * Unicode normalization NFD * Get rid of BOM

cmake : fix typo in AMDGPU_TARGETS (ggerganov#7356)

ef277de

ggml : fix quants nans when all the group weights are very close to z…

0583484

…ero (ggerganov#7313)

convert : fix set_vocab_sentencepiece (ggerganov#6866)

b49a13d

* convert : fix set_vocab_sentencepiece * Update convert-hf-to-gguf.py

github-actions-labeler: initial commit (ggerganov#7330)

de73196

* github-actions-labeler: initial commit [no ci] * github actions: remove priority auto labeling [no ci]

Update and fix Vulkan soft_max and argsort implementations (ggerganov…

c1b295e

…#7237) * Update and fix Vulkan softmax implementation * Update and fix Vulkan argsort implementation

perplexity : ndot progress and show stats with < 100 tasks (ggerganov…

ca57e0f

…#7348) Fix floating point error with ndot printing, allow end stats on lower task numbers if multiple-choice tasks.

cuda : add half2 __shfl_xor() for ROCm 5.5 (ggerganov#7263)

d233b50

server: correct --threads documentation [no ci] (ggerganov#7362)

cb42c29

CUDA: deduplicate FlashAttention code (ggerganov#7352)

133d99c

android : use "ci-android" branch for CI (ggerganov#7341)

511182e

* android : use "ci-android" branch for CI * ggml : disable SIMD exp and silu for 32-bit ARM ggml-ci * android : do not fetch, use add_subdirectory instead * cmake : provide binary dir

ci : re-enable sanitizer runs (ggerganov#7358)

059031b

* Revert "ci : temporary disable sanitizer builds (ggerganov#6128)" This reverts commit 4f6d133. * ci : trigger

cmake : update android comments (ggerganov#7341)

854d365

cuda : clear error after buffer allocation failure (ggerganov#7376)

ab33f7a

aahouzi and others added 29 commits May 19, 2024 22:46

Add StableLM2 pre-tokenizer (ggerganov#7349)

6aade19

* Add StableLM pre-tokenizer * Fix space * Fix trailing whitespace

server: fix seed being reported back (ggerganov#7382)

4185839

server: add test for token probs (ggerganov#7347)

1b01f06

ggml: implement quantized KV cache for FA (ggerganov#7372)

5ca49cb

ggml : fix another case of quants nans (ggerganov#7387)

e4e6f67

quantize : fix --keep-split check (ggerganov#7374)

1ea2a00

llama : remove MPI backend (ggerganov#7395)

d359f30

Add provisions for windows support for BF16 code including CMake prov…

33c8d50

…ision for enabling AVX512_BF16 (ggerganov#7258)

tests : fix --keep_split -> --keep-split (ggerganov#7374)

2789baf

server : return error on too large embedding input (ggerganov#7389)

e932094

server : tuning tests (ggerganov#7388)

1cc0155

* server : don't pass temperature as string * server : increase timeout * tests : fix the fix 0.8f -> 0.8 ggml-ci * tests : set explicit temperature

ggml-opencl, llama: using reserve() if count already known (ggerganov…

213e90e

…#7272)

Update README.md (ggerganov#7410)

26cd423

[SYCL] Update SYCL upscale operation (ggerganov#7321)

6bf9b66

* Update SYCL upscale operation * Formatting * Remove messages

server : fix temperature + disable some tests (ggerganov#7409)

3bc10cb

* server : fix temperature * server : disable tests relying on parallel determinism * ci : change server Debug -> RelWithDebInfo

rpc : track allocated buffers (ggerganov#7411)

db10f01

* rpc : track allocated buffers ref: ggerganov#7407 * rpc : pack rpc_tensor tightly

perplexity: update README FP16 results [no ci] (ggerganov#7413)

20385ce

llama : remove Persimmon (ggerganov#7408)

fabf30b

* llama : remove Persimmon * requirements : remove

CUDA: deduplicate mmq code (ggerganov#7397)

d8ee902

tests : test-tokenizer-0.sh print more info (ggerganov#7402)

c3f8d58

CUDA: fix unused warning in mmq.cu (ggerganov#7442)

fcf6538

grammars: fix resampling logic regression (ggerganov#7424)

e402de3

metal : handle F16 inf values, fix FA partial offload (ggerganov#7434)

6369bf0

ggml-ci

l3utterfly merged commit ece1194 into layla-build May 22, 2024
66 of 80 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

merge from upstream #20

merge from upstream #20

l3utterfly commented May 22, 2024

merge from upstream #20

merge from upstream #20

Conversation

l3utterfly commented May 22, 2024