Merge branch 'master' into flipflop-stream

Add '--flipflop-offload' startup argument
2026-02-16 21:21:05 +00:00 · 2025-10-28 15:08:27 -07:00 · 2025-10-13 21:10:44 -07:00 · 2025-10-13 21:04:37 -07:00 · 2025-10-03 16:21:01 -07:00 · 2025-10-03 14:32:56 -07:00
126 changed files with 6260 additions and 10522 deletions
--- a/.ci/windows_amd_base_files/README_VERY_IMPORTANT.txt
+++ b/.ci/windows_amd_base_files/README_VERY_IMPORTANT.txt
@@ -1,5 +1,5 @@
-As of the time of writing this you need this driver for best results:
-https://www.amd.com/en/resources/support-articles/release-notes/RN-AMDGPU-WINDOWS-PYTORCH-7-1-1.html
+As of the time of writing this you need this preview driver for best results:
+https://www.amd.com/en/resources/support-articles/release-notes/RN-AMDGPU-WINDOWS-PYTORCH-PREVIEW.html

 HOW TO RUN:

@@ -25,4 +25,3 @@ In the ComfyUI directory you will find a file: extra_model_paths.yaml.example
 Rename this file to: extra_model_paths.yaml and edit it with your favorite text editor.


-
--- a/.github/ISSUE_TEMPLATE/bug-report.yml
+++ b/.github/ISSUE_TEMPLATE/bug-report.yml
@@ -8,15 +8,13 @@ body:
        Before submitting a **Bug Report**, please ensure the following:

        - **1:** You are running the latest version of ComfyUI.
-        - **2:** You have your ComfyUI logs and relevant workflow on hand and will post them in this bug report.
+        - **2:** You have looked at the existing bug reports and made sure this isn't already reported.
        - **3:** You confirmed that the bug is not caused by a custom node. You can disable all custom nodes by passing
-        `--disable-all-custom-nodes` command line argument. If you have custom node try updating them to the latest version.
+        `--disable-all-custom-nodes` command line argument.
        - **4:** This is an actual bug in ComfyUI, not just a support question. A bug is when you can specify exact
        steps to replicate what went wrong and others will be able to repeat your steps and see the same issue happen.

-        ## Very Important
-
-        Please make sure that you post ALL your ComfyUI logs in the bug report. A bug report without logs will likely be ignored.
+        If unsure, ask on the [ComfyUI Matrix Space](https://app.element.io/#/room/%23comfyui_space%3Amatrix.org) or the [Comfy Org Discord](https://discord.gg/comfyorg) first.
  - type: checkboxes
    id: custom-nodes-test
    attributes:
--- a/.github/PULL_REQUEST_TEMPLATE/api-node.md
+++ b/.github/PULL_REQUEST_TEMPLATE/api-node.md
@@ -1,21 +0,0 @@
-<!-- API_NODE_PR_CHECKLIST: do not remove -->
-
-## API Node PR Checklist
-
-### Scope
- [ ] **Is API Node Change**
-
-### Pricing & Billing
- [ ] **Need pricing update**
- [ ] **No pricing update**
-
-If **Need pricing update**:
- [ ] Metronome rate cards updated
- [ ] Auto‑billing tests updated and passing
-
-### QA
- [ ] **QA done**
- [ ] **QA not required**
-
-### Comms
- [ ] Informed **Kosinkadink**
--- a/.github/workflows/api-node-template.yml
+++ b/.github/workflows/api-node-template.yml
@@ -1,58 +0,0 @@
-name: Append API Node PR template
-
-on:
-  pull_request_target:
-    types: [opened, reopened, synchronize, ready_for_review]
-    paths:
-      - 'comfy_api_nodes/**'   # only run if these files changed
-
-permissions:
-  contents: read
-  pull-requests: write
-
-jobs:
-  inject:
-    runs-on: ubuntu-latest
-    steps:
-      - name: Ensure template exists and append to PR body
-        uses: actions/github-script@v7
-        with:
-          script: |
-            const { owner, repo } = context.repo;
-            const number = context.payload.pull_request.number;
-            const templatePath = '.github/PULL_REQUEST_TEMPLATE/api-node.md';
-            const marker = '<!-- API_NODE_PR_CHECKLIST: do not remove -->';
-
-            const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: number });
-
-            let templateText;
-            try {
-              const res = await github.rest.repos.getContent({
-                owner,
-                repo,
-                path: templatePath,
-                ref: pr.base.ref
-              });
-              const buf = Buffer.from(res.data.content, res.data.encoding || 'base64');
-              templateText = buf.toString('utf8');
-            } catch (e) {
-              core.setFailed(`Required PR template not found at "${templatePath}" on ${pr.base.ref}. Please add it to the repo.`);
-              return;
-            }
-
-            // Enforce the presence of the marker inside the template (for idempotence)
-            if (!templateText.includes(marker)) {
-              core.setFailed(`Template at "${templatePath}" does not contain the required marker:\n${marker}\nAdd it so we can detect duplicates safely.`);
-              return;
-            }
-
-            // If the PR already contains the marker, do not append again.
-            const body = pr.body || '';
-            if (body.includes(marker)) {
-              core.info('Template already present in PR body; nothing to inject.');
-              return;
-            }
-
-            const newBody = (body ? body + '\n\n' : '') + templateText + '\n';
-            await github.rest.pulls.update({ owner, repo, pull_number: number, body: newBody });
-            core.notice('API Node template appended to PR description.');
--- a/.github/workflows/release-stable-all.yml
+++ b/.github/workflows/release-stable-all.yml
@@ -14,7 +14,7 @@ jobs:
      contents: "write"
      packages: "write"
      pull-requests: "read"
-    name: "Release NVIDIA Default (cu130)"
+    name: "Release NVIDIA Default (cu129)"
    uses: ./.github/workflows/stable-release.yml
    with:
      git_tag: ${{ inputs.git_tag }}
@@ -43,33 +43,16 @@ jobs:
      test_release: true
    secrets: inherit

-  release_nvidia_cu126:
-    permissions:
-      contents: "write"
-      packages: "write"
-      pull-requests: "read"
-    name: "Release NVIDIA cu126"
-    uses: ./.github/workflows/stable-release.yml
-    with:
-      git_tag: ${{ inputs.git_tag }}
-      cache_tag: "cu126"
-      python_minor: "12"
-      python_patch: "10"
-      rel_name: "nvidia"
-      rel_extra_name: "_cu126"
-      test_release: true
-    secrets: inherit
-
  release_amd_rocm:
    permissions:
      contents: "write"
      packages: "write"
      pull-requests: "read"
-    name: "Release AMD ROCm 7.1.1"
+    name: "Release AMD ROCm 6.4.4"
    uses: ./.github/workflows/stable-release.yml
    with:
      git_tag: ${{ inputs.git_tag }}
-      cache_tag: "rocm711"
+      cache_tag: "rocm644"
      python_minor: "12"
      python_patch: "10"
      rel_name: "amd"
--- a/.github/workflows/test-ci.yml
+++ b/.github/workflows/test-ci.yml
@@ -21,15 +21,14 @@ jobs:
      fail-fast: false
      matrix:
        # os: [macos, linux, windows]
-        # os: [macos, linux]
-        os: [linux]
-        python_version: ["3.10", "3.11", "3.12"]
+        os: [macos, linux]
+        python_version: ["3.9", "3.10", "3.11", "3.12"]
        cuda_version: ["12.1"]
        torch_version: ["stable"]
        include:
-          # - os: macos
-          #   runner_label: [self-hosted, macOS]
-          #   flags: "--use-pytorch-cross-attention"
+          - os: macos
+            runner_label: [self-hosted, macOS]
+            flags: "--use-pytorch-cross-attention"
          - os: linux
            runner_label: [self-hosted, Linux]
            flags: ""
@@ -74,15 +73,14 @@ jobs:
    strategy:
      fail-fast: false
      matrix:
-        # os: [macos, linux]
-        os: [linux]
+        os: [macos, linux]
        python_version: ["3.11"]
        cuda_version: ["12.1"]
        torch_version: ["nightly"]
        include:
-          # - os: macos
-          #   runner_label: [self-hosted, macOS]
-          #   flags: "--use-pytorch-cross-attention"
+          - os: macos
+            runner_label: [self-hosted, macOS]
+            flags: "--use-pytorch-cross-attention"
          - os: linux
            runner_label: [self-hosted, Linux]
            flags: ""
--- a/QUANTIZATION.md
+++ b/QUANTIZATION.md
@@ -1,168 +0,0 @@
-# The Comfy guide to Quantization
-
-
-## How does quantization work?
-
-Quantization aims to map a high-precision value x_f to a lower precision format with minimal loss in accuracy. These smaller formats then serve to reduce the models memory footprint and increase throughput by using specialized hardware.
-
-When simply converting a value from FP16 to FP8 using the round-nearest method we might hit two issues:
- The dynamic range of FP16 (-65,504, 65,504) far exceeds FP8 formats like E4M3 (-448, 448) or E5M2 (-57,344, 57,344), potentially resulting in clipped values
- The original values are concentrated in a small range (e.g. -1,1) leaving many FP8-bits "unused"
-
-By using a scaling factor, we aim to map these values into the quantized-dtype range, making use of the full spectrum. One of the easiest approaches, and common, is using per-tensor absolute-maximum scaling.
-
-```
-absmax = max(abs(tensor))
-scale = amax / max_dynamic_range_low_precision
-
-# Quantization
-tensor_q = (tensor / scale).to(low_precision_dtype)
-
-# De-Quantization
-tensor_dq = tensor_q.to(fp16) * scale
-
-tensor_dq ~ tensor
-```
-
-Given that additional information (scaling factor) is needed to "interpret" the quantized values, we describe those as derived datatypes.
-
-
-## Quantization in Comfy
-
-```
-QuantizedTensor (torch.Tensor subclass)
-  ↓ __torch_dispatch__
-Two-Level Registry (generic + layout handlers)
-  ↓
-MixedPrecisionOps + Metadata Detection
-```
-
-### Representation
-
-To represent these derived datatypes, ComfyUI uses a subclass of torch.Tensor to implements these using the `QuantizedTensor` class found in `comfy/quant_ops.py`
-
-A `Layout` class defines how a specific quantization format behaves:
- Required parameters
- Quantize method
- De-Quantize method
-
-```python
-from comfy.quant_ops import QuantizedLayout
-
-class MyLayout(QuantizedLayout):
-    @classmethod
-    def quantize(cls, tensor, **kwargs):
-        # Convert to quantized format
-        qdata = ...
-        params = {'scale': ..., 'orig_dtype': tensor.dtype}
-        return qdata, params
-    
-    @staticmethod
-    def dequantize(qdata, scale, orig_dtype, **kwargs):
-        return qdata.to(orig_dtype) * scale
-```
-
-To then run operations using these QuantizedTensors we use two registry systems to define supported operations. 
-The first is a **generic registry** that handles operations common to all quantized formats (e.g., `.to()`, `.clone()`, `.reshape()`).
-
-The second registry is layout-specific and allows to implement fast-paths like nn.Linear.
-```python
-from comfy.quant_ops import register_layout_op
-
-@register_layout_op(torch.ops.aten.linear.default, MyLayout)
-def my_linear(func, args, kwargs):
-    # Extract tensors, call optimized kernel
-    ...
-```
-When `torch.nn.functional.linear()` is called with QuantizedTensor arguments, `__torch_dispatch__` automatically routes to the registered implementation.
-For any unsupported operation, QuantizedTensor will fallback to call `dequantize` and dispatch using the high-precision implementation.
-
-
-### Mixed Precision
-
-The `MixedPrecisionOps` class (lines 542-648 in `comfy/ops.py`) enables per-layer quantization decisions, allowing different layers in a model to use different precisions. This is activated when a model config contains a `layer_quant_config` dictionary that specifies which layers should be quantized and how.
-
-**Architecture:**
-
-```python
-class MixedPrecisionOps(disable_weight_init):
-    _layer_quant_config = {}  # Maps layer names to quantization configs
-    _compute_dtype = torch.bfloat16  # Default compute / dequantize precision
-```
-
-**Key mechanism:**
-
-The custom `Linear._load_from_state_dict()` method inspects each layer during model loading:
- If the layer name is **not** in `_layer_quant_config`: load weight as regular tensor in `_compute_dtype`
- If the layer name **is** in `_layer_quant_config`: 
-  - Load weight as `QuantizedTensor` with the specified layout (e.g., `TensorCoreFP8Layout`)
-  - Load associated quantization parameters (scales, block_size, etc.)
-
-**Why it's needed:**
-
-Not all layers tolerate quantization equally. Sensitive operations like final projections can be kept in higher precision, while compute-heavy matmuls are quantized. This provides most of the performance benefits while maintaining quality.
-
-The system is selected in `pick_operations()` when `model_config.layer_quant_config` is present, making it the highest-priority operation mode.
-
-
-## Checkpoint Format
-
-Quantized checkpoints are stored as standard safetensors files with quantized weight tensors and associated scaling parameters, plus a `_quantization_metadata` JSON entry describing the quantization scheme.
-
-The quantized checkpoint will contain the same layers as the original checkpoint but:
- The weights are stored as quantized values, sometimes using a different storage datatype. E.g. uint8 container for fp8.
- For each quantized weight a number of additional scaling parameters are stored alongside depending on the recipe.
- We store a metadata.json in the metadata of the final safetensor containing the `_quantization_metadata` describing which layers are quantized and what layout has been used.
-
-### Scaling Parameters details
-We define 4 possible scaling parameters that should cover most recipes in the near-future:
- **weight_scale**: quantization scalers for the weights
- **weight_scale_2**: global scalers in the context of double scaling
- **pre_quant_scale**: scalers used for smoothing salient weights
- **input_scale**: quantization scalers for the activations
-
-| Format | Storage dtype | weight_scale | weight_scale_2 | pre_quant_scale | input_scale |
-|--------|---------------|--------------|----------------|-----------------|-------------|
-| float8_e4m3fn | float32 | float32 (scalar) | - | - | float32 (scalar) |
-
-You can find the defined formats in `comfy/quant_ops.py` (QUANT_ALGOS).
-
-### Quantization Metadata
-
-The metadata stored alongside the checkpoint contains:
- **format_version**: String to define a version of the standard
- **layers**: A dictionary mapping layer names to their quantization format. The format string maps to the definitions found in `QUANT_ALGOS`. 
-
-Example:
-```json
-{
-  "_quantization_metadata": {
-    "format_version": "1.0",
-    "layers": {
-      "model.layers.0.mlp.up_proj": "float8_e4m3fn",
-      "model.layers.0.mlp.down_proj": "float8_e4m3fn",
-      "model.layers.1.mlp.up_proj": "float8_e4m3fn"
-    }
-  }
-}
-```
-
-
-## Creating Quantized Checkpoints
-
-To create compatible checkpoints, use any quantization tool provided the output follows the checkpoint format described above and uses a layout defined in `QUANT_ALGOS`.
-
-### Weight Quantization
-
-Weight quantization is straightforward - compute the scaling factor directly from the weight tensor using the absolute maximum method described earlier. Each layer's weights are quantized independently and stored with their corresponding `weight_scale` parameter.
-
-### Calibration (for Activation Quantization)
-
-Activation quantization (e.g., for FP8 Tensor Core operations) requires `input_scale` parameters that cannot be determined from static weights alone. Since activation values depend on actual inputs, we use **post-training calibration (PTQ)**:
-
-1. **Collect statistics**: Run inference on N representative samples
-2. **Track activations**: Record the absolute maximum (`amax`) of inputs to each quantized layer
-3. **Compute scales**: Derive `input_scale` from collected statistics
-4. **Store in checkpoint**: Save `input_scale` parameters alongside weights
-
-The calibration dataset should be representative of your target use case. For diffusion models, this typically means a diverse set of prompts and generation parameters.
--- a/README.md
+++ b/README.md
@@ -67,8 +67,6 @@ See what ComfyUI can do with the [example workflows](https://comfyanonymous.gith
   - [HiDream](https://comfyanonymous.github.io/ComfyUI_examples/hidream/)
   - [Qwen Image](https://comfyanonymous.github.io/ComfyUI_examples/qwen_image/)
   - [Hunyuan Image 2.1](https://comfyanonymous.github.io/ComfyUI_examples/hunyuan_image/)
-   - [Flux 2](https://comfyanonymous.github.io/ComfyUI_examples/flux2/)
-   - [Z Image](https://comfyanonymous.github.io/ComfyUI_examples/z_image/)
 - Image Editing Models
   - [Omnigen 2](https://comfyanonymous.github.io/ComfyUI_examples/omnigen/)
   - [Flux Kontext](https://comfyanonymous.github.io/ComfyUI_examples/flux/#flux-kontext-image-editing-model)
@@ -114,11 +112,10 @@ Workflow examples can be found on the [Examples page](https://comfyanonymous.git

 ## Release Process

-ComfyUI follows a weekly release cycle targeting Monday but this regularly changes because of model releases or large changes to the codebase. There are three interconnected repositories:
+ComfyUI follows a weekly release cycle targeting Friday but this regularly changes because of model releases or large changes to the codebase. There are three interconnected repositories:

 1. **[ComfyUI Core](https://github.com/comfyanonymous/ComfyUI)**
-   - Releases a new stable version (e.g., v0.7.0) roughly every week.
-   - Commits outside of the stable release tags may be very unstable and break many custom nodes.
+   - Releases a new stable version (e.g., v0.7.0)
   - Serves as the foundation for the desktop release

 2. **[ComfyUI Desktop](https://github.com/Comfy-Org/desktop)**
@@ -175,7 +172,7 @@ There is a portable standalone build for Windows that should work for running on

 ### [Direct link to download](https://github.com/comfyanonymous/ComfyUI/releases/latest/download/ComfyUI_windows_portable_nvidia.7z)

-Simply download, extract with [7-Zip](https://7-zip.org) or with the windows explorer on recent windows versions and run. For smaller models you normally only need to put the checkpoints (the huge ckpt/safetensors files) in: ComfyUI\models\checkpoints but many of the larger models have multiple files. Make sure to follow the instructions to know which subfolder to put them in ComfyUI\models\
+Simply download, extract with [7-Zip](https://7-zip.org) and run. Make sure you put your Stable Diffusion checkpoints/models (the huge ckpt/safetensors files) in: ComfyUI\models\checkpoints

 If you have trouble extracting it, right click the file -> properties -> unblock

@@ -185,9 +182,7 @@ Update your Nvidia drivers if it doesn't start.

 [Experimental portable for AMD GPUs](https://github.com/comfyanonymous/ComfyUI/releases/latest/download/ComfyUI_windows_portable_amd.7z)

-[Portable with pytorch cuda 12.8 and python 3.12](https://github.com/comfyanonymous/ComfyUI/releases/latest/download/ComfyUI_windows_portable_nvidia_cu128.7z).
-
-[Portable with pytorch cuda 12.6 and python 3.12](https://github.com/comfyanonymous/ComfyUI/releases/latest/download/ComfyUI_windows_portable_nvidia_cu126.7z) (Supports Nvidia 10 series and older GPUs).
+[Portable with pytorch cuda 12.8 and python 3.12](https://github.com/comfyanonymous/ComfyUI/releases/latest/download/ComfyUI_windows_portable_nvidia_cu128.7z) (Supports Nvidia 10 series and older GPUs).

 #### How do I share models between another UI and ComfyUI?

@@ -204,7 +199,7 @@ comfy install

 ## Manual Install (Windows, Linux)

-Python 3.14 works but you may encounter issues with the torch compile node. The free threaded variant is still missing some dependencies.
+Python 3.14 will work if you comment out the `kornia` dependency in the requirements.txt file (breaks the canny node) but it is not recommended.

 Python 3.13 is very well supported. If you have trouble with some custom node dependencies on 3.13 you can try 3.12

@@ -225,7 +220,7 @@ AMD users can install rocm and pytorch with pip if you don't have it already ins

 This is the command to install the nightly with ROCm 7.0 which might have some performance improvements:

-```pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/rocm7.1```
+```pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/rocm7.0```


 ### AMD GPUs (Experimental: Windows and Linux), RDNA 3, 3.5 and 4 only.
@@ -246,7 +241,7 @@ RDNA 4 (RX 9000 series):

 ### Intel GPUs (Windows and Linux)

-Intel Arc GPU users can install native PyTorch with torch.xpu support using pip. More information can be found [here](https://pytorch.org/docs/main/notes/get_start_xpu.html)
+(Option 1) Intel Arc GPU users can install native PyTorch with torch.xpu support using pip. More information can be found [here](https://pytorch.org/docs/main/notes/get_start_xpu.html)

 1. To install PyTorch xpu, use the following command:

@@ -256,6 +251,10 @@ This is the command to install the Pytorch xpu nightly which might have some per

 ```pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/xpu```

+(Option 2) Alternatively, Intel GPUs supported by Intel Extension for PyTorch (IPEX) can leverage IPEX for improved performance.
+
+1. visit [Installation](https://intel.github.io/intel-extension-for-pytorch/index.html#installation?platform=gpu) for more information.
+
 ### NVIDIA

 Nvidia users should install stable pytorch using this command:
--- a/app/frontend_management.py
+++ b/app/frontend_management.py
@@ -10,8 +10,7 @@ import importlib
 from dataclasses import dataclass
 from functools import cached_property
 from pathlib import Path
-from typing import Dict, TypedDict, Optional
-from aiohttp import web
+from typing import TypedDict, Optional
 from importlib.metadata import version

 import requests
@@ -258,54 +257,7 @@ comfyui-frontend-package is not installed.
            sys.exit(-1)

    @classmethod
-    def template_asset_map(cls) -> Optional[Dict[str, str]]:
-        """Return a mapping of template asset names to their absolute paths."""
-        try:
-            from comfyui_workflow_templates import (
-                get_asset_path,
-                iter_templates,
-            )
-        except ImportError:
-            logging.error(
-                f"""
-********** ERROR ***********
-
-comfyui-workflow-templates is not installed.
-
-{frontend_install_warning_message()}
-
-********** ERROR ***********
-""".strip()
-            )
-            return None
-
-        try:
-            template_entries = list(iter_templates())
-        except Exception as exc:
-            logging.error(f"Failed to enumerate workflow templates: {exc}")
-            return None
-
-        asset_map: Dict[str, str] = {}
-        try:
-            for entry in template_entries:
-                for asset in entry.assets:
-                    asset_map[asset.filename] = get_asset_path(
-                        entry.template_id, asset.filename
-                    )
-        except Exception as exc:
-            logging.error(f"Failed to resolve template asset paths: {exc}")
-            return None
-
-        if not asset_map:
-            logging.error("No workflow template assets found. Did the packages install correctly?")
-            return None
-
-        return asset_map
-
-
-    @classmethod
-    def legacy_templates_path(cls) -> Optional[str]:
-        """Return the legacy templates directory shipped inside the meta package."""
+    def templates_path(cls) -> str:
        try:
            import comfyui_workflow_templates

@@ -324,7 +276,6 @@ comfyui-workflow-templates is not installed.
 ********** ERROR ***********
 """.strip()
            )
-            return None

    @classmethod
    def embedded_docs_path(cls) -> str:
@@ -441,17 +392,3 @@ comfyui-workflow-templates is not installed.
            logging.info("Falling back to the default frontend.")
            check_frontend_version()
            return cls.default_frontend_path()
-    @classmethod
-    def template_asset_handler(cls):
-        assets = cls.template_asset_map()
-        if not assets:
-            return None
-
-        async def serve_template(request: web.Request) -> web.StreamResponse:
-            rel_path = request.match_info.get("path", "")
-            target = assets.get(rel_path)
-            if target is None:
-                raise web.HTTPNotFound()
-            return web.FileResponse(target)
-
-        return serve_template
--- a/app/user_manager.py
+++ b/app/user_manager.py
@@ -59,9 +59,6 @@ class UserManager():
        user = "default"
        if args.multi_user and "comfy-user" in request.headers:
            user = request.headers["comfy-user"]
-            # Block System Users (use same error message to prevent probing)
-            if user.startswith(folder_paths.SYSTEM_USER_PREFIX):
-                raise KeyError("Unknown user: " + user)

        if user not in self.users:
            raise KeyError("Unknown user: " + user)
@@ -69,16 +66,15 @@ class UserManager():
        return user

    def get_request_user_filepath(self, request, file, type="userdata", create_dir=True):
+        user_directory = folder_paths.get_user_directory()
+
        if type == "userdata":
-            root_dir = folder_paths.get_user_directory()
+            root_dir = user_directory
        else:
            raise KeyError("Unknown filepath type:" + type)

        user = self.get_request_user_id(request)
-        user_root = folder_paths.get_public_user_directory(user)
-        if user_root is None:
-            return None
-        path = user_root
+        path = user_root = os.path.abspath(os.path.join(root_dir, user))

        # prevent leaving /{type}
        if os.path.commonpath((root_dir, user_root)) != root_dir:
@@ -105,11 +101,7 @@ class UserManager():
        name = name.strip()
        if not name:
            raise ValueError("username not provided")
-        if name.startswith(folder_paths.SYSTEM_USER_PREFIX):
-            raise ValueError("System User prefix not allowed")
        user_id = re.sub("[^a-zA-Z0-9-_]+", '-', name)
-        if user_id.startswith(folder_paths.SYSTEM_USER_PREFIX):
-            raise ValueError("System User prefix not allowed")
        user_id = user_id + "_" + str(uuid.uuid4())

        self.users[user_id] = name
@@ -140,10 +132,7 @@ class UserManager():
            if username in self.users.values():
                return web.json_response({"error": "Duplicate username."}, status=400)

-            try:
-                user_id = self.add_user(username)
-            except ValueError as e:
-                return web.json_response({"error": str(e)}, status=400)
+            user_id = self.add_user(username)
            return web.json_response(user_id)

        @routes.get("/userdata")
@@ -435,7 +424,7 @@ class UserManager():
                return source

            dest = get_user_data_path(request, check_exists=False, param="dest")
-            if not isinstance(dest, str):
+            if not isinstance(source, str):
                return dest

            overwrite = request.query.get("overwrite", 'true') != "false"
--- a/comfy/cldm/cldm.py
+++ b/comfy/cldm/cldm.py
@@ -413,8 +413,7 @@ class ControlNet(nn.Module):
        out_middle = []

        if self.num_classes is not None:
-            if y is None:
-                raise ValueError("y is None, did you try using a controlnet for SDXL on SD1?")
+            assert y.shape[0] == x.shape[0]
            emb = emb + self.label_emb(y)

        h = x
--- a/comfy/cli_args.py
+++ b/comfy/cli_args.py
@@ -105,7 +105,6 @@ cache_group = parser.add_mutually_exclusive_group()
 cache_group.add_argument("--cache-classic", action="store_true", help="Use the old style (aggressive) caching.")
 cache_group.add_argument("--cache-lru", type=int, default=0, help="Use LRU caching with a maximum of N node results cached. May use more RAM/VRAM.")
 cache_group.add_argument("--cache-none", action="store_true", help="Reduced RAM/VRAM usage at the expense of executing every node for each run.")
-cache_group.add_argument("--cache-ram", nargs='?', const=4.0, type=float, default=0, help="Use RAM pressure caching with the specified headroom threshold. If available RAM drops below the threhold the cache remove large items to free RAM. Default 4GB")

 attn_group = parser.add_mutually_exclusive_group()
 attn_group.add_argument("--use-split-cross-attention", action="store_true", help="Use the split cross attention optimization. Ignored when xformers is used.")
@@ -131,8 +130,9 @@ vram_group.add_argument("--cpu", action="store_true", help="To use the CPU for e

 parser.add_argument("--reserve-vram", type=float, default=None, help="Set the amount of vram in GB you want to reserve for use by your OS/other software. By default some amount is reserved depending on your OS.")

-parser.add_argument("--async-offload", nargs='?', const=2, type=int, default=None, metavar="NUM_STREAMS", help="Use async weight offloading. An optional argument controls the amount of offload streams. Default is 2. Enabled by default on Nvidia.")
-parser.add_argument("--disable-async-offload", action="store_true", help="Disable async weight offloading.")
+parser.add_argument("--async-offload", action="store_true", help="Use async weight offloading.")
+
+parser.add_argument("--flipflop-offload", action="store_true", help="Use async flipflop weight offloading for supported DiT models.")

 parser.add_argument("--force-non-blocking", action="store_true", help="Force ComfyUI to use non-blocking operations for all applicable tensors. This may improve performance on some non-Nvidia systems but can cause issues with some workflows.")

@@ -147,9 +147,7 @@ class PerformanceFeature(enum.Enum):
    CublasOps = "cublas_ops"
    AutoTune = "autotune"

-parser.add_argument("--fast", nargs="*", type=PerformanceFeature, help="Enable some untested and potentially quality deteriorating optimizations. This is used to test new features so using it might crash your comfyui. --fast with no arguments enables everything. You can pass a list specific optimizations if you only want to enable specific ones. Current valid optimizations: {}".format(" ".join(map(lambda c: c.value, PerformanceFeature))))
-
-parser.add_argument("--disable-pinned-memory", action="store_true", help="Disable pinned memory use.")
+parser.add_argument("--fast", nargs="*", type=PerformanceFeature, help="Enable some untested and potentially quality deteriorating optimizations. --fast with no arguments enables everything. You can pass a list specific optimizations if you only want to enable specific ones. Current valid optimizations: {}".format(" ".join(map(lambda c: c.value, PerformanceFeature))))

 parser.add_argument("--mmap-torch-files", action="store_true", help="Use mmap when loading ckpt/pt files.")
 parser.add_argument("--disable-mmap", action="store_true", help="Don't use mmap when loading safetensors.")
@@ -161,7 +159,7 @@ parser.add_argument("--windows-standalone-build", action="store_true", help="Win
 parser.add_argument("--disable-metadata", action="store_true", help="Disable saving prompt metadata in files.")
 parser.add_argument("--disable-all-custom-nodes", action="store_true", help="Disable loading all custom nodes.")
 parser.add_argument("--whitelist-custom-nodes", type=str, nargs='+', default=[], help="Specify custom node folders to load even when --disable-all-custom-nodes is enabled.")
-parser.add_argument("--disable-api-nodes", action="store_true", help="Disable loading all api nodes. Also prevents the frontend from communicating with the internet.")
+parser.add_argument("--disable-api-nodes", action="store_true", help="Disable loading all api nodes.")

 parser.add_argument("--multi-user", action="store_true", help="Enables per-user storage.")

--- a/comfy/controlnet.py
+++ b/comfy/controlnet.py
@@ -310,13 +310,11 @@ class ControlLoraOps:
            self.bias = None

        def forward(self, input):
-            weight, bias, offload_stream = comfy.ops.cast_bias_weight(self, input, offloadable=True)
+            weight, bias = comfy.ops.cast_bias_weight(self, input)
            if self.up is not None:
-                x = torch.nn.functional.linear(input, weight + (torch.mm(self.up.flatten(start_dim=1), self.down.flatten(start_dim=1))).reshape(self.weight.shape).type(input.dtype), bias)
+                return torch.nn.functional.linear(input, weight + (torch.mm(self.up.flatten(start_dim=1), self.down.flatten(start_dim=1))).reshape(self.weight.shape).type(input.dtype), bias)
            else:
-                x = torch.nn.functional.linear(input, weight, bias)
-            comfy.ops.uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+                return torch.nn.functional.linear(input, weight, bias)

    class Conv2d(torch.nn.Module, comfy.ops.CastWeightBiasOp):
        def __init__(
@@ -352,13 +350,12 @@ class ControlLoraOps:


        def forward(self, input):
-            weight, bias, offload_stream = comfy.ops.cast_bias_weight(self, input, offloadable=True)
+            weight, bias = comfy.ops.cast_bias_weight(self, input)
            if self.up is not None:
-                x = torch.nn.functional.conv2d(input, weight + (torch.mm(self.up.flatten(start_dim=1), self.down.flatten(start_dim=1))).reshape(self.weight.shape).type(input.dtype), bias, self.stride, self.padding, self.dilation, self.groups)
+                return torch.nn.functional.conv2d(input, weight + (torch.mm(self.up.flatten(start_dim=1), self.down.flatten(start_dim=1))).reshape(self.weight.shape).type(input.dtype), bias, self.stride, self.padding, self.dilation, self.groups)
            else:
-                x = torch.nn.functional.conv2d(input, weight, bias, self.stride, self.padding, self.dilation, self.groups)
-            comfy.ops.uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+                return torch.nn.functional.conv2d(input, weight, bias, self.stride, self.padding, self.dilation, self.groups)
+

 class ControlLora(ControlNet):
    def __init__(self, control_weights, global_average_pooling=False, model_options={}): #TODO? model_options
--- a/comfy/latent_formats.py
+++ b/comfy/latent_formats.py
@@ -6,7 +6,6 @@ class LatentFormat:
    latent_dimensions = 2
    latent_rgb_factors = None
    latent_rgb_factors_bias = None
-    latent_rgb_factors_reshape = None
    taesd_decoder_name = None

    def process_in(self, latent):
@@ -179,54 +178,6 @@ class Flux(SD3):
    def process_out(self, latent):
        return (latent / self.scale_factor) + self.shift_factor

-class Flux2(LatentFormat):
-    latent_channels = 128
-
-    def __init__(self):
-        self.latent_rgb_factors =[
-            [0.0058, 0.0113, 0.0073],
-            [0.0495, 0.0443, 0.0836],
-            [-0.0099, 0.0096, 0.0644],
-            [0.2144, 0.3009, 0.3652],
-            [0.0166, -0.0039, -0.0054],
-            [0.0157, 0.0103, -0.0160],
-            [-0.0398, 0.0902, -0.0235],
-            [-0.0052, 0.0095, 0.0109],
-            [-0.3527, -0.2712, -0.1666],
-            [-0.0301, -0.0356, -0.0180],
-            [-0.0107, 0.0078, 0.0013],
-            [0.0746, 0.0090, -0.0941],
-            [0.0156, 0.0169, 0.0070],
-            [-0.0034, -0.0040, -0.0114],
-            [0.0032, 0.0181, 0.0080],
-            [-0.0939, -0.0008, 0.0186],
-            [0.0018, 0.0043, 0.0104],
-            [0.0284, 0.0056, -0.0127],
-            [-0.0024, -0.0022, -0.0030],
-            [0.1207, -0.0026, 0.0065],
-            [0.0128, 0.0101, 0.0142],
-            [0.0137, -0.0072, -0.0007],
-            [0.0095, 0.0092, -0.0059],
-            [0.0000, -0.0077, -0.0049],
-            [-0.0465, -0.0204, -0.0312],
-            [0.0095, 0.0012, -0.0066],
-            [0.0290, -0.0034, 0.0025],
-            [0.0220, 0.0169, -0.0048],
-            [-0.0332, -0.0457, -0.0468],
-            [-0.0085, 0.0389, 0.0609],
-            [-0.0076, 0.0003, -0.0043],
-            [-0.0111, -0.0460, -0.0614],
-        ]
-
-        self.latent_rgb_factors_bias = [-0.0329, -0.0718, -0.0851]
-        self.latent_rgb_factors_reshape = lambda t: t.reshape(t.shape[0], 32, 2, 2, t.shape[-2], t.shape[-1]).permute(0, 1, 4, 2, 5, 3).reshape(t.shape[0], 32, t.shape[-2] * 2, t.shape[-1] * 2)
-
-    def process_in(self, latent):
-        return latent
-
-    def process_out(self, latent):
-        return latent
-
 class Mochi(LatentFormat):
    latent_channels = 12
    latent_dimensions = 3
@@ -431,7 +382,6 @@ class HunyuanVideo(LatentFormat):
    ]

    latent_rgb_factors_bias = [ 0.0259, -0.0192, -0.0761]
-    taesd_decoder_name = "taehv"

 class Cosmos1CV8x8x8(LatentFormat):
    latent_channels = 16
@@ -495,7 +445,7 @@ class Wan21(LatentFormat):
        ]).view(1, self.latent_channels, 1, 1, 1)


-        self.taesd_decoder_name = "lighttaew2_1"
+        self.taesd_decoder_name = None #TODO

    def process_in(self, latent):
        latents_mean = self.latents_mean.to(latent.device, latent.dtype)
@@ -566,7 +516,6 @@ class Wan22(Wan21):

    def __init__(self):
        self.scale_factor = 1.0
-        self.taesd_decoder_name = "lighttaew2_2"
        self.latents_mean = torch.tensor([
                -0.2289, -0.0052, -0.1323, -0.2339, -0.2799, 0.0174, 0.1838, 0.1557,
                -0.1382, 0.0542, 0.2813, 0.0891, 0.1570, -0.0098, 0.0375, -0.1825,
@@ -662,67 +611,6 @@ class HunyuanImage21Refiner(LatentFormat):
    latent_dimensions = 3
    scale_factor = 1.03682

-    def process_in(self, latent):
-        out = latent * self.scale_factor
-        out = torch.cat((out[:, :, :1], out), dim=2)
-        out = out.permute(0, 2, 1, 3, 4)
-        b, f_times_2, c, h, w = out.shape
-        out = out.reshape(b, f_times_2 // 2, 2 * c, h, w)
-        out = out.permute(0, 2, 1, 3, 4).contiguous()
-        return out
-
-    def process_out(self, latent):
-        z = latent / self.scale_factor
-        z = z.permute(0, 2, 1, 3, 4)
-        b, f, c, h, w = z.shape
-        z = z.reshape(b, f, 2, c // 2, h, w)
-        z = z.permute(0, 1, 2, 3, 4, 5).reshape(b, f * 2, c // 2, h, w)
-        z = z.permute(0, 2, 1, 3, 4)
-        z = z[:, :, 1:]
-        return z
-
-class HunyuanVideo15(LatentFormat):
-    latent_rgb_factors = [
-        [ 0.0568, -0.0521, -0.0131],
-        [ 0.0014,  0.0735,  0.0326],
-        [ 0.0186,  0.0531, -0.0138],
-        [-0.0031,  0.0051,  0.0288],
-        [ 0.0110,  0.0556,  0.0432],
-        [-0.0041, -0.0023, -0.0485],
-        [ 0.0530,  0.0413,  0.0253],
-        [ 0.0283,  0.0251,  0.0339],
-        [ 0.0277, -0.0372, -0.0093],
-        [ 0.0393,  0.0944,  0.1131],
-        [ 0.0020,  0.0251,  0.0037],
-        [-0.0017,  0.0012,  0.0234],
-        [ 0.0468,  0.0436,  0.0203],
-        [ 0.0354,  0.0439, -0.0233],
-        [ 0.0090,  0.0123,  0.0346],
-        [ 0.0382,  0.0029,  0.0217],
-        [ 0.0261, -0.0300,  0.0030],
-        [-0.0088, -0.0220, -0.0283],
-        [-0.0272, -0.0121, -0.0363],
-        [-0.0664, -0.0622,  0.0144],
-        [ 0.0414,  0.0479,  0.0529],
-        [ 0.0355,  0.0612, -0.0247],
-        [ 0.0147,  0.0264,  0.0174],
-        [ 0.0438,  0.0038,  0.0542],
-        [ 0.0431, -0.0573, -0.0033],
-        [-0.0162, -0.0211, -0.0406],
-        [-0.0487, -0.0295, -0.0393],
-        [ 0.0005, -0.0109,  0.0253],
-        [ 0.0296,  0.0591,  0.0353],
-        [ 0.0119,  0.0181, -0.0306],
-        [-0.0085, -0.0362,  0.0229],
-        [ 0.0005, -0.0106,  0.0242]
-    ]
-
-    latent_rgb_factors_bias = [ 0.0456, -0.0202, -0.0644]
-    latent_channels = 32
-    latent_dimensions = 3
-    scale_factor = 1.03682
-    taesd_decoder_name = "lighttaehy1_5"
-
 class Hunyuan3Dv2(LatentFormat):
    latent_channels = 64
    latent_dimensions = 1
--- a/comfy/ldm/chroma/layers.py
+++ b/comfy/ldm/chroma/layers.py
@@ -1,15 +1,15 @@
 import torch
 from torch import Tensor, nn

+from comfy.ldm.flux.math import attention
 from comfy.ldm.flux.layers import (
    MLPEmbedder,
    RMSNorm,
+    QKNorm,
+    SelfAttention,
    ModulationOut,
 )

-# TODO: remove this in a few months
-SingleStreamBlock = None
-DoubleStreamBlock = None


 class ChromaModulationOut(ModulationOut):
@@ -48,6 +48,124 @@ class Approximator(nn.Module):
        return x


+class DoubleStreamBlock(nn.Module):
+    def __init__(self, hidden_size: int, num_heads: int, mlp_ratio: float, qkv_bias: bool = False, flipped_img_txt=False, dtype=None, device=None, operations=None):
+        super().__init__()
+
+        mlp_hidden_dim = int(hidden_size * mlp_ratio)
+        self.num_heads = num_heads
+        self.hidden_size = hidden_size
+        self.img_norm1 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
+        self.img_attn = SelfAttention(dim=hidden_size, num_heads=num_heads, qkv_bias=qkv_bias, dtype=dtype, device=device, operations=operations)
+
+        self.img_norm2 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
+        self.img_mlp = nn.Sequential(
+            operations.Linear(hidden_size, mlp_hidden_dim, bias=True, dtype=dtype, device=device),
+            nn.GELU(approximate="tanh"),
+            operations.Linear(mlp_hidden_dim, hidden_size, bias=True, dtype=dtype, device=device),
+        )
+
+        self.txt_norm1 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
+        self.txt_attn = SelfAttention(dim=hidden_size, num_heads=num_heads, qkv_bias=qkv_bias, dtype=dtype, device=device, operations=operations)
+
+        self.txt_norm2 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
+        self.txt_mlp = nn.Sequential(
+            operations.Linear(hidden_size, mlp_hidden_dim, bias=True, dtype=dtype, device=device),
+            nn.GELU(approximate="tanh"),
+            operations.Linear(mlp_hidden_dim, hidden_size, bias=True, dtype=dtype, device=device),
+        )
+        self.flipped_img_txt = flipped_img_txt
+
+    def forward(self, img: Tensor, txt: Tensor, pe: Tensor, vec: Tensor, attn_mask=None, transformer_options={}):
+        (img_mod1, img_mod2), (txt_mod1, txt_mod2) = vec
+
+        # prepare image for attention
+        img_modulated = torch.addcmul(img_mod1.shift, 1 + img_mod1.scale, self.img_norm1(img))
+        img_qkv = self.img_attn.qkv(img_modulated)
+        img_q, img_k, img_v = img_qkv.view(img_qkv.shape[0], img_qkv.shape[1], 3, self.num_heads, -1).permute(2, 0, 3, 1, 4)
+        img_q, img_k = self.img_attn.norm(img_q, img_k, img_v)
+
+        # prepare txt for attention
+        txt_modulated = torch.addcmul(txt_mod1.shift, 1 + txt_mod1.scale, self.txt_norm1(txt))
+        txt_qkv = self.txt_attn.qkv(txt_modulated)
+        txt_q, txt_k, txt_v = txt_qkv.view(txt_qkv.shape[0], txt_qkv.shape[1], 3, self.num_heads, -1).permute(2, 0, 3, 1, 4)
+        txt_q, txt_k = self.txt_attn.norm(txt_q, txt_k, txt_v)
+
+        # run actual attention
+        attn = attention(torch.cat((txt_q, img_q), dim=2),
+                         torch.cat((txt_k, img_k), dim=2),
+                         torch.cat((txt_v, img_v), dim=2),
+                         pe=pe, mask=attn_mask, transformer_options=transformer_options)
+
+        txt_attn, img_attn = attn[:, : txt.shape[1]], attn[:, txt.shape[1] :]
+
+        # calculate the img bloks
+        img.addcmul_(img_mod1.gate, self.img_attn.proj(img_attn))
+        img.addcmul_(img_mod2.gate, self.img_mlp(torch.addcmul(img_mod2.shift, 1 + img_mod2.scale, self.img_norm2(img))))
+
+        # calculate the txt bloks
+        txt.addcmul_(txt_mod1.gate, self.txt_attn.proj(txt_attn))
+        txt.addcmul_(txt_mod2.gate, self.txt_mlp(torch.addcmul(txt_mod2.shift, 1 + txt_mod2.scale, self.txt_norm2(txt))))
+
+        if txt.dtype == torch.float16:
+            txt = torch.nan_to_num(txt, nan=0.0, posinf=65504, neginf=-65504)
+
+        return img, txt
+
+
+class SingleStreamBlock(nn.Module):
+    """
+    A DiT block with parallel linear layers as described in
+    https://arxiv.org/abs/2302.05442 and adapted modulation interface.
+    """
+
+    def __init__(
+        self,
+        hidden_size: int,
+        num_heads: int,
+        mlp_ratio: float = 4.0,
+        qk_scale: float = None,
+        dtype=None,
+        device=None,
+        operations=None
+    ):
+        super().__init__()
+        self.hidden_dim = hidden_size
+        self.num_heads = num_heads
+        head_dim = hidden_size // num_heads
+        self.scale = qk_scale or head_dim**-0.5
+
+        self.mlp_hidden_dim = int(hidden_size * mlp_ratio)
+        # qkv and mlp_in
+        self.linear1 = operations.Linear(hidden_size, hidden_size * 3 + self.mlp_hidden_dim, dtype=dtype, device=device)
+        # proj and mlp_out
+        self.linear2 = operations.Linear(hidden_size + self.mlp_hidden_dim, hidden_size, dtype=dtype, device=device)
+
+        self.norm = QKNorm(head_dim, dtype=dtype, device=device, operations=operations)
+
+        self.hidden_size = hidden_size
+        self.pre_norm = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
+
+        self.mlp_act = nn.GELU(approximate="tanh")
+
+    def forward(self, x: Tensor, pe: Tensor, vec: Tensor, attn_mask=None, transformer_options={}) -> Tensor:
+        mod = vec
+        x_mod = torch.addcmul(mod.shift, 1 + mod.scale, self.pre_norm(x))
+        qkv, mlp = torch.split(self.linear1(x_mod), [3 * self.hidden_size, self.mlp_hidden_dim], dim=-1)
+
+        q, k, v = qkv.view(qkv.shape[0], qkv.shape[1], 3, self.num_heads, -1).permute(2, 0, 3, 1, 4)
+        q, k = self.norm(q, k, v)
+
+        # compute attention
+        attn = attention(q, k, v, pe=pe, mask=attn_mask, transformer_options=transformer_options)
+        # compute activation in mlp stream, cat again and run second linear layer
+        output = self.linear2(torch.cat((attn, self.mlp_act(mlp)), 2))
+        x.addcmul_(mod.gate, output)
+        if x.dtype == torch.float16:
+            x = torch.nan_to_num(x, nan=0.0, posinf=65504, neginf=-65504)
+        return x
+
+
 class LastLayer(nn.Module):
    def __init__(self, hidden_size: int, patch_size: int, out_channels: int, dtype=None, device=None, operations=None):
        super().__init__()
--- a/comfy/ldm/chroma/model.py
+++ b/comfy/ldm/chroma/model.py
@@ -11,12 +11,12 @@ import comfy.ldm.common_dit
 from comfy.ldm.flux.layers import (
    EmbedND,
    timestep_embedding,
-    DoubleStreamBlock,
-    SingleStreamBlock,
 )

 from .layers import (
+    DoubleStreamBlock,
    LastLayer,
+    SingleStreamBlock,
    Approximator,
    ChromaModulationOut,
 )
@@ -90,7 +90,6 @@ class Chroma(nn.Module):
                    self.num_heads,
                    mlp_ratio=params.mlp_ratio,
                    qkv_bias=params.qkv_bias,
-                    modulation=False,
                    dtype=dtype, device=device, operations=operations
                )
                for _ in range(params.depth)
@@ -99,7 +98,7 @@ class Chroma(nn.Module):

        self.single_blocks = nn.ModuleList(
            [
-                SingleStreamBlock(self.hidden_size, self.num_heads, mlp_ratio=params.mlp_ratio, modulation=False, dtype=dtype, device=device, operations=operations)
+                SingleStreamBlock(self.hidden_size, self.num_heads, mlp_ratio=params.mlp_ratio, dtype=dtype, device=device, operations=operations)
                for _ in range(params.depth_single_blocks)
            ]
        )
@@ -179,10 +178,7 @@ class Chroma(nn.Module):
        pe = self.pe_embedder(ids)

        blocks_replace = patches_replace.get("dit", {})
-        transformer_options["total_blocks"] = len(self.double_blocks)
-        transformer_options["block_type"] = "double"
        for i, block in enumerate(self.double_blocks):
-            transformer_options["block_index"] = i
            if i not in self.skip_mmdit:
                double_mod = (
                    self.get_modulations(mod_vectors, "double_img", idx=i),
@@ -225,10 +221,7 @@ class Chroma(nn.Module):

        img = torch.cat((txt, img), 1)

-        transformer_options["total_blocks"] = len(self.single_blocks)
-        transformer_options["block_type"] = "single"
        for i, block in enumerate(self.single_blocks):
-            transformer_options["block_index"] = i
            if i not in self.skip_dit:
                single_mod = self.get_modulations(mod_vectors, "single", idx=i)
                if ("single_block", i) in blocks_replace:
--- a/comfy/ldm/chroma_radiance/model.py
+++ b/comfy/ldm/chroma_radiance/model.py
@@ -10,10 +10,12 @@ from torch import Tensor, nn
 from einops import repeat
 import comfy.ldm.common_dit

-from comfy.ldm.flux.layers import EmbedND, DoubleStreamBlock, SingleStreamBlock
+from comfy.ldm.flux.layers import EmbedND

 from comfy.ldm.chroma.model import Chroma, ChromaParams
 from comfy.ldm.chroma.layers import (
+    DoubleStreamBlock,
+    SingleStreamBlock,
    Approximator,
 )
 from .layers import (
@@ -87,6 +89,7 @@ class ChromaRadiance(Chroma):
                    dtype=dtype, device=device, operations=operations
                )

+
        self.double_blocks = nn.ModuleList(
            [
                DoubleStreamBlock(
@@ -94,7 +97,6 @@ class ChromaRadiance(Chroma):
                    self.num_heads,
                    mlp_ratio=params.mlp_ratio,
                    qkv_bias=params.qkv_bias,
-                    modulation=False,
                    dtype=dtype, device=device, operations=operations
                )
                for _ in range(params.depth)
@@ -107,7 +109,6 @@ class ChromaRadiance(Chroma):
                    self.hidden_size,
                    self.num_heads,
                    mlp_ratio=params.mlp_ratio,
-                    modulation=False,
                    dtype=dtype, device=device, operations=operations,
                )
                for _ in range(params.depth_single_blocks)
--- a/comfy/ldm/flipflop_transformer.py
+++ b/comfy/ldm/flipflop_transformer.py
@@ -0,0 +1,200 @@
+from __future__ import annotations
+import torch
+import copy
+
+import comfy.model_management
+
+
+class FlipFlopModule(torch.nn.Module):
+    def __init__(self, block_types: tuple[str, ...], enable_flipflop: bool = True):
+        super().__init__()
+        self.block_types = block_types
+        self.enable_flipflop = enable_flipflop
+        self.flipflop: dict[str, FlipFlopHolder] = {}
+        self.block_info: dict[str, tuple[int, int]] = {}
+        self.flipflop_prefixes: list[str] = []
+
+    def setup_flipflop_holders(self, block_info: dict[str, tuple[int, int]], flipflop_prefixes: list[str], load_device: torch.device, offload_device: torch.device):
+        for block_type, (flipflop_blocks, total_blocks) in block_info.items():
+            if block_type in self.flipflop:
+                continue
+            self.flipflop[block_type] = FlipFlopHolder(getattr(self, block_type)[total_blocks-flipflop_blocks:], flipflop_blocks, total_blocks, load_device, offload_device)
+            self.block_info[block_type] = (flipflop_blocks, total_blocks)
+        self.flipflop_prefixes = flipflop_prefixes.copy()
+
+    def init_flipflop_block_copies(self, device: torch.device) -> int:
+        memory_freed = 0
+        for holder in self.flipflop.values():
+            memory_freed += holder.init_flipflop_block_copies(device)
+        return memory_freed
+
+    def clean_flipflop_holders(self):
+        memory_freed = 0
+        for block_type in list(self.flipflop.keys()):
+            memory_freed += self.flipflop[block_type].clean_flipflop_blocks()
+            del self.flipflop[block_type]
+        self.block_info = {}
+        self.flipflop_prefixes = []
+        return memory_freed
+
+    def get_all_blocks(self, block_type: str) -> list[torch.nn.Module]:
+        return getattr(self, block_type)
+
+    def get_blocks(self, block_type: str) -> torch.nn.ModuleList:
+        if block_type not in self.block_types:
+            raise ValueError(f"Block type {block_type} not found in {self.block_types}")
+        if block_type in self.flipflop:
+            return getattr(self, block_type)[:self.flipflop[block_type].i_offset]
+        return getattr(self, block_type)
+
+    def get_all_block_module_sizes(self, reverse_sort_by_size: bool = False) -> list[tuple[str, int]]:
+        '''
+        Returns a list of (block_type, size) sorted by size.
+        If reverse_sort_by_size is True, the list is sorted by size in reverse order.
+        '''
+        sizes = [(block_type, self.get_block_module_size(block_type)) for block_type in self.block_types]
+        sizes.sort(key=lambda x: x[1], reverse=reverse_sort_by_size)
+        return sizes
+
+    def get_block_module_size(self, block_type: str) -> int:
+        return comfy.model_management.module_size(getattr(self, block_type)[0])
+
+    def execute_blocks(self, block_type: str, func, out: torch.Tensor | tuple[torch.Tensor,...], *args, **kwargs):
+        # execute blocks, supporting both single and double (or higher) block types
+        if isinstance(out, torch.Tensor):
+            out = (out,)
+        for i, block in enumerate(self.get_blocks(block_type)):
+            out = func(i, block, *out, *args, **kwargs)
+            if isinstance(out, torch.Tensor):
+                out = (out,)
+        if block_type in self.flipflop:
+            holder = self.flipflop[block_type]
+            with holder.context() as ctx:
+                for i, block in enumerate(holder.blocks):
+                    out = ctx(func, i, block, *out, *args, **kwargs)
+                    if isinstance(out, torch.Tensor):
+                        out = (out,)
+        if len(out) == 1:
+            out = out[0]
+        return out
+
+
+class FlipFlopContext:
+    def __init__(self, holder: FlipFlopHolder):
+        # NOTE: there is a bug when there are an odd number of blocks to flipflop.
+        # Worked around right now by always making sure it will be even, but need to resolve.
+        self.holder = holder
+        self.reset()
+
+    def reset(self):
+        self.num_blocks = len(self.holder.blocks)
+        self.first_flip = True
+        self.first_flop = True
+        self.last_flip = False
+        self.last_flop = False
+
+    def __enter__(self):
+        self.reset()
+        return self
+
+    def __exit__(self, exc_type, exc_value, traceback):
+        self.holder.compute_stream.record_event(self.holder.cpy_end_event)
+
+    def do_flip(self, func, i: int, _, *args, **kwargs):
+        # flip
+        self.holder.compute_stream.wait_event(self.holder.cpy_end_event)
+        with torch.cuda.stream(self.holder.compute_stream):
+            out = func(i+self.holder.i_offset, self.holder.flip, *args, **kwargs)
+        self.holder.event_flip.record(self.holder.compute_stream)
+        # while flip executes, queue flop to copy to its next block
+        next_flop_i = i + 1
+        if next_flop_i >= self.num_blocks:
+            next_flop_i = next_flop_i - self.num_blocks
+            self.last_flip = True
+        if not self.first_flip:
+            self.holder._copy_state_dict(self.holder.flop.state_dict(), self.holder.blocks[next_flop_i].state_dict(), self.holder.event_flop, self.holder.cpy_end_event)
+        if self.last_flip:
+            self.holder._copy_state_dict(self.holder.flip.state_dict(), self.holder.blocks[0].state_dict(), cpy_start_event=self.holder.event_flip)
+        self.first_flip = False
+        return out
+
+    def do_flop(self, func, i: int, _, *args, **kwargs):
+        # flop
+        if not self.first_flop:
+            self.holder.compute_stream.wait_event(self.holder.cpy_end_event)
+        with torch.cuda.stream(self.holder.compute_stream):
+            out = func(i+self.holder.i_offset, self.holder.flop, *args, **kwargs)
+        self.holder.event_flop.record(self.holder.compute_stream)
+        # while flop executes, queue flip to copy to its next block
+        next_flip_i = i + 1
+        if next_flip_i >= self.num_blocks:
+            next_flip_i = next_flip_i - self.num_blocks
+            self.last_flop = True
+        self.holder._copy_state_dict(self.holder.flip.state_dict(), self.holder.blocks[next_flip_i].state_dict(), self.holder.event_flip, self.holder.cpy_end_event)
+        if self.last_flop:
+            self.holder._copy_state_dict(self.holder.flop.state_dict(), self.holder.blocks[1].state_dict(), cpy_start_event=self.holder.event_flop)
+        self.first_flop = False
+        return out
+
+    @torch.no_grad()
+    def __call__(self, func, i: int, block: torch.nn.Module, *args, **kwargs):
+        # flips are even indexes, flops are odd indexes
+        if i % 2 == 0:
+            return self.do_flip(func, i, block, *args, **kwargs)
+        else:
+            return self.do_flop(func, i, block, *args, **kwargs)
+
+
+class FlipFlopHolder:
+    def __init__(self, blocks: list[torch.nn.Module], flip_amount: int, total_amount: int, load_device: torch.device, offload_device: torch.device):
+        self.load_device = load_device
+        self.offload_device = offload_device
+        self.blocks = blocks
+        self.flip_amount = flip_amount
+        self.total_amount = total_amount
+        # NOTE: used to make sure block indexes passed into block functions match expected patch indexes
+        self.i_offset = total_amount - flip_amount
+
+        self.block_module_size = 0
+        if len(self.blocks) > 0:
+            self.block_module_size = comfy.model_management.module_size(self.blocks[0])
+
+        self.flip: torch.nn.Module = None
+        self.flop: torch.nn.Module = None
+
+        self.compute_stream = torch.cuda.default_stream(self.load_device)
+        self.cpy_stream = torch.cuda.Stream(self.load_device)
+
+        self.event_flip = torch.cuda.Event(enable_timing=False)
+        self.event_flop = torch.cuda.Event(enable_timing=False)
+        self.cpy_end_event = torch.cuda.Event(enable_timing=False)
+        # INIT - is this actually needed?
+        self.compute_stream.record_event(self.cpy_end_event)
+
+    def _copy_state_dict(self, dst, src, cpy_start_event: torch.cuda.Event=None, cpy_end_event: torch.cuda.Event=None):
+        if cpy_start_event:
+            self.cpy_stream.wait_event(cpy_start_event)
+
+        with torch.cuda.stream(self.cpy_stream):
+            for k, v in src.items():
+                dst[k].copy_(v, non_blocking=True)
+        if cpy_end_event:
+            cpy_end_event.record(self.cpy_stream)
+
+    def context(self):
+        return FlipFlopContext(self)
+
+    def init_flipflop_block_copies(self, load_device: torch.device) -> int:
+        self.flip = copy.deepcopy(self.blocks[0]).to(device=load_device)
+        self.flop = copy.deepcopy(self.blocks[1]).to(device=load_device)
+        return comfy.model_management.module_size(self.flip) + comfy.model_management.module_size(self.flop)
+
+    def clean_flipflop_blocks(self) -> int:
+        memory_freed = 0
+        memory_freed += comfy.model_management.module_size(self.flip)
+        memory_freed += comfy.model_management.module_size(self.flop)
+        del self.flip
+        del self.flop
+        self.flip = None
+        self.flop = None
+        return memory_freed
--- a/comfy/ldm/flux/layers.py
+++ b/comfy/ldm/flux/layers.py
@@ -48,11 +48,11 @@ def timestep_embedding(t: Tensor, dim, max_period=10000, time_factor: float = 10
    return embedding

 class MLPEmbedder(nn.Module):
-    def __init__(self, in_dim: int, hidden_dim: int, bias=True, dtype=None, device=None, operations=None):
+    def __init__(self, in_dim: int, hidden_dim: int, dtype=None, device=None, operations=None):
        super().__init__()
-        self.in_layer = operations.Linear(in_dim, hidden_dim, bias=bias, dtype=dtype, device=device)
+        self.in_layer = operations.Linear(in_dim, hidden_dim, bias=True, dtype=dtype, device=device)
        self.silu = nn.SiLU()
-        self.out_layer = operations.Linear(hidden_dim, hidden_dim, bias=bias, dtype=dtype, device=device)
+        self.out_layer = operations.Linear(hidden_dim, hidden_dim, bias=True, dtype=dtype, device=device)

    def forward(self, x: Tensor) -> Tensor:
        return self.out_layer(self.silu(self.in_layer(x)))
@@ -80,14 +80,14 @@ class QKNorm(torch.nn.Module):


 class SelfAttention(nn.Module):
-    def __init__(self, dim: int, num_heads: int = 8, qkv_bias: bool = False, proj_bias: bool = True, dtype=None, device=None, operations=None):
+    def __init__(self, dim: int, num_heads: int = 8, qkv_bias: bool = False, dtype=None, device=None, operations=None):
        super().__init__()
        self.num_heads = num_heads
        head_dim = dim // num_heads

        self.qkv = operations.Linear(dim, dim * 3, bias=qkv_bias, dtype=dtype, device=device)
        self.norm = QKNorm(head_dim, dtype=dtype, device=device, operations=operations)
-        self.proj = operations.Linear(dim, dim, bias=proj_bias, dtype=dtype, device=device)
+        self.proj = operations.Linear(dim, dim, dtype=dtype, device=device)


@dataclass
@@ -98,11 +98,11 @@ class ModulationOut:


 class Modulation(nn.Module):
-    def __init__(self, dim: int, double: bool, bias=True, dtype=None, device=None, operations=None):
+    def __init__(self, dim: int, double: bool, dtype=None, device=None, operations=None):
        super().__init__()
        self.is_double = double
        self.multiplier = 6 if double else 3
-        self.lin = operations.Linear(dim, self.multiplier * dim, bias=bias, dtype=dtype, device=device)
+        self.lin = operations.Linear(dim, self.multiplier * dim, bias=True, dtype=dtype, device=device)

    def forward(self, vec: Tensor) -> tuple:
        if vec.ndim == 2:
@@ -129,129 +129,77 @@ def apply_mod(tensor, m_mult, m_add=None, modulation_dims=None):
        return tensor


-class SiLUActivation(nn.Module):
-    def __init__(self):
-        super().__init__()
-        self.gate_fn = nn.SiLU()
-
-    def forward(self, x: Tensor) -> Tensor:
-        x1, x2 = x.chunk(2, dim=-1)
-        return self.gate_fn(x1) * x2
-
-
 class DoubleStreamBlock(nn.Module):
-    def __init__(self, hidden_size: int, num_heads: int, mlp_ratio: float, qkv_bias: bool = False, flipped_img_txt=False, modulation=True, mlp_silu_act=False, proj_bias=True, dtype=None, device=None, operations=None):
+    def __init__(self, hidden_size: int, num_heads: int, mlp_ratio: float, qkv_bias: bool = False, flipped_img_txt=False, dtype=None, device=None, operations=None):
        super().__init__()

        mlp_hidden_dim = int(hidden_size * mlp_ratio)
        self.num_heads = num_heads
        self.hidden_size = hidden_size
-        self.modulation = modulation
-
-        if self.modulation:
-            self.img_mod = Modulation(hidden_size, double=True, dtype=dtype, device=device, operations=operations)
-
+        self.img_mod = Modulation(hidden_size, double=True, dtype=dtype, device=device, operations=operations)
        self.img_norm1 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
-        self.img_attn = SelfAttention(dim=hidden_size, num_heads=num_heads, qkv_bias=qkv_bias, proj_bias=proj_bias, dtype=dtype, device=device, operations=operations)
+        self.img_attn = SelfAttention(dim=hidden_size, num_heads=num_heads, qkv_bias=qkv_bias, dtype=dtype, device=device, operations=operations)

        self.img_norm2 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
+        self.img_mlp = nn.Sequential(
+            operations.Linear(hidden_size, mlp_hidden_dim, bias=True, dtype=dtype, device=device),
+            nn.GELU(approximate="tanh"),
+            operations.Linear(mlp_hidden_dim, hidden_size, bias=True, dtype=dtype, device=device),
+        )

-        if mlp_silu_act:
-            self.img_mlp = nn.Sequential(
-                operations.Linear(hidden_size, mlp_hidden_dim * 2, bias=False, dtype=dtype, device=device),
-                SiLUActivation(),
-                operations.Linear(mlp_hidden_dim, hidden_size, bias=False, dtype=dtype, device=device),
-            )
-        else:
-            self.img_mlp = nn.Sequential(
-                operations.Linear(hidden_size, mlp_hidden_dim, bias=True, dtype=dtype, device=device),
-                nn.GELU(approximate="tanh"),
-                operations.Linear(mlp_hidden_dim, hidden_size, bias=True, dtype=dtype, device=device),
-            )
-
-        if self.modulation:
-            self.txt_mod = Modulation(hidden_size, double=True, dtype=dtype, device=device, operations=operations)
-
+        self.txt_mod = Modulation(hidden_size, double=True, dtype=dtype, device=device, operations=operations)
        self.txt_norm1 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
-        self.txt_attn = SelfAttention(dim=hidden_size, num_heads=num_heads, qkv_bias=qkv_bias, proj_bias=proj_bias, dtype=dtype, device=device, operations=operations)
+        self.txt_attn = SelfAttention(dim=hidden_size, num_heads=num_heads, qkv_bias=qkv_bias, dtype=dtype, device=device, operations=operations)

        self.txt_norm2 = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
-
-        if mlp_silu_act:
-            self.txt_mlp = nn.Sequential(
-                operations.Linear(hidden_size, mlp_hidden_dim * 2, bias=False, dtype=dtype, device=device),
-                SiLUActivation(),
-                operations.Linear(mlp_hidden_dim, hidden_size, bias=False, dtype=dtype, device=device),
-            )
-        else:
-            self.txt_mlp = nn.Sequential(
-                operations.Linear(hidden_size, mlp_hidden_dim, bias=True, dtype=dtype, device=device),
-                nn.GELU(approximate="tanh"),
-                operations.Linear(mlp_hidden_dim, hidden_size, bias=True, dtype=dtype, device=device),
-            )
-
+        self.txt_mlp = nn.Sequential(
+            operations.Linear(hidden_size, mlp_hidden_dim, bias=True, dtype=dtype, device=device),
+            nn.GELU(approximate="tanh"),
+            operations.Linear(mlp_hidden_dim, hidden_size, bias=True, dtype=dtype, device=device),
+        )
        self.flipped_img_txt = flipped_img_txt

    def forward(self, img: Tensor, txt: Tensor, vec: Tensor, pe: Tensor, attn_mask=None, modulation_dims_img=None, modulation_dims_txt=None, transformer_options={}):
-        if self.modulation:
-            img_mod1, img_mod2 = self.img_mod(vec)
-            txt_mod1, txt_mod2 = self.txt_mod(vec)
-        else:
-            (img_mod1, img_mod2), (txt_mod1, txt_mod2) = vec
+        img_mod1, img_mod2 = self.img_mod(vec)
+        txt_mod1, txt_mod2 = self.txt_mod(vec)

        # prepare image for attention
        img_modulated = self.img_norm1(img)
        img_modulated = apply_mod(img_modulated, (1 + img_mod1.scale), img_mod1.shift, modulation_dims_img)
        img_qkv = self.img_attn.qkv(img_modulated)
-        del img_modulated
        img_q, img_k, img_v = img_qkv.view(img_qkv.shape[0], img_qkv.shape[1], 3, self.num_heads, -1).permute(2, 0, 3, 1, 4)
-        del img_qkv
        img_q, img_k = self.img_attn.norm(img_q, img_k, img_v)

        # prepare txt for attention
        txt_modulated = self.txt_norm1(txt)
        txt_modulated = apply_mod(txt_modulated, (1 + txt_mod1.scale), txt_mod1.shift, modulation_dims_txt)
        txt_qkv = self.txt_attn.qkv(txt_modulated)
-        del txt_modulated
        txt_q, txt_k, txt_v = txt_qkv.view(txt_qkv.shape[0], txt_qkv.shape[1], 3, self.num_heads, -1).permute(2, 0, 3, 1, 4)
-        del txt_qkv
        txt_q, txt_k = self.txt_attn.norm(txt_q, txt_k, txt_v)

        if self.flipped_img_txt:
-            q = torch.cat((img_q, txt_q), dim=2)
-            del img_q, txt_q
-            k = torch.cat((img_k, txt_k), dim=2)
-            del img_k, txt_k
-            v = torch.cat((img_v, txt_v), dim=2)
-            del img_v, txt_v
            # run actual attention
-            attn = attention(q, k, v,
+            attn = attention(torch.cat((img_q, txt_q), dim=2),
+                             torch.cat((img_k, txt_k), dim=2),
+                             torch.cat((img_v, txt_v), dim=2),
                             pe=pe, mask=attn_mask, transformer_options=transformer_options)
-            del q, k, v

            img_attn, txt_attn = attn[:, : img.shape[1]], attn[:, img.shape[1]:]
        else:
-            q = torch.cat((txt_q, img_q), dim=2)
-            del txt_q, img_q
-            k = torch.cat((txt_k, img_k), dim=2)
-            del txt_k, img_k
-            v = torch.cat((txt_v, img_v), dim=2)
-            del txt_v, img_v
            # run actual attention
-            attn = attention(q, k, v,
+            attn = attention(torch.cat((txt_q, img_q), dim=2),
+                             torch.cat((txt_k, img_k), dim=2),
+                             torch.cat((txt_v, img_v), dim=2),
                             pe=pe, mask=attn_mask, transformer_options=transformer_options)
-            del q, k, v

            txt_attn, img_attn = attn[:, : txt.shape[1]], attn[:, txt.shape[1]:]

        # calculate the img bloks
-        img += apply_mod(self.img_attn.proj(img_attn), img_mod1.gate, None, modulation_dims_img)
-        del img_attn
-        img += apply_mod(self.img_mlp(apply_mod(self.img_norm2(img), (1 + img_mod2.scale), img_mod2.shift, modulation_dims_img)), img_mod2.gate, None, modulation_dims_img)
+        img = img + apply_mod(self.img_attn.proj(img_attn), img_mod1.gate, None, modulation_dims_img)
+        img = img + apply_mod(self.img_mlp(apply_mod(self.img_norm2(img), (1 + img_mod2.scale), img_mod2.shift, modulation_dims_img)), img_mod2.gate, None, modulation_dims_img)

        # calculate the txt bloks
        txt += apply_mod(self.txt_attn.proj(txt_attn), txt_mod1.gate, None, modulation_dims_txt)
-        del txt_attn
        txt += apply_mod(self.txt_mlp(apply_mod(self.txt_norm2(txt), (1 + txt_mod2.scale), txt_mod2.shift, modulation_dims_txt)), txt_mod2.gate, None, modulation_dims_txt)

        if txt.dtype == torch.float16:
@@ -272,9 +220,6 @@ class SingleStreamBlock(nn.Module):
        num_heads: int,
        mlp_ratio: float = 4.0,
        qk_scale: float = None,
-        modulation=True,
-        mlp_silu_act=False,
-        bias=True,
        dtype=None,
        device=None,
        operations=None
@@ -286,47 +231,30 @@ class SingleStreamBlock(nn.Module):
        self.scale = qk_scale or head_dim**-0.5

        self.mlp_hidden_dim = int(hidden_size * mlp_ratio)
-
-        self.mlp_hidden_dim_first = self.mlp_hidden_dim
-        if mlp_silu_act:
-            self.mlp_hidden_dim_first = int(hidden_size * mlp_ratio * 2)
-            self.mlp_act = SiLUActivation()
-        else:
-            self.mlp_act = nn.GELU(approximate="tanh")
-
        # qkv and mlp_in
-        self.linear1 = operations.Linear(hidden_size, hidden_size * 3 + self.mlp_hidden_dim_first, bias=bias, dtype=dtype, device=device)
+        self.linear1 = operations.Linear(hidden_size, hidden_size * 3 + self.mlp_hidden_dim, dtype=dtype, device=device)
        # proj and mlp_out
-        self.linear2 = operations.Linear(hidden_size + self.mlp_hidden_dim, hidden_size, bias=bias, dtype=dtype, device=device)
+        self.linear2 = operations.Linear(hidden_size + self.mlp_hidden_dim, hidden_size, dtype=dtype, device=device)

        self.norm = QKNorm(head_dim, dtype=dtype, device=device, operations=operations)

        self.hidden_size = hidden_size
        self.pre_norm = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)

-        if modulation:
-            self.modulation = Modulation(hidden_size, double=False, dtype=dtype, device=device, operations=operations)
-        else:
-            self.modulation = None
+        self.mlp_act = nn.GELU(approximate="tanh")
+        self.modulation = Modulation(hidden_size, double=False, dtype=dtype, device=device, operations=operations)

    def forward(self, x: Tensor, vec: Tensor, pe: Tensor, attn_mask=None, modulation_dims=None, transformer_options={}) -> Tensor:
-        if self.modulation:
-            mod, _ = self.modulation(vec)
-        else:
-            mod = vec
-
-        qkv, mlp = torch.split(self.linear1(apply_mod(self.pre_norm(x), (1 + mod.scale), mod.shift, modulation_dims)), [3 * self.hidden_size, self.mlp_hidden_dim_first], dim=-1)
+        mod, _ = self.modulation(vec)
+        qkv, mlp = torch.split(self.linear1(apply_mod(self.pre_norm(x), (1 + mod.scale), mod.shift, modulation_dims)), [3 * self.hidden_size, self.mlp_hidden_dim], dim=-1)

        q, k, v = qkv.view(qkv.shape[0], qkv.shape[1], 3, self.num_heads, -1).permute(2, 0, 3, 1, 4)
-        del qkv
        q, k = self.norm(q, k, v)

        # compute attention
        attn = attention(q, k, v, pe=pe, mask=attn_mask, transformer_options=transformer_options)
-        del q, k, v
        # compute activation in mlp stream, cat again and run second linear layer
-        mlp = self.mlp_act(mlp)
-        output = self.linear2(torch.cat((attn, mlp), 2))
+        output = self.linear2(torch.cat((attn, self.mlp_act(mlp)), 2))
        x += apply_mod(output, mod.gate, None, modulation_dims)
        if x.dtype == torch.float16:
            x = torch.nan_to_num(x, nan=0.0, posinf=65504, neginf=-65504)
@@ -334,11 +262,11 @@ class SingleStreamBlock(nn.Module):


 class LastLayer(nn.Module):
-    def __init__(self, hidden_size: int, patch_size: int, out_channels: int, bias=True, dtype=None, device=None, operations=None):
+    def __init__(self, hidden_size: int, patch_size: int, out_channels: int, dtype=None, device=None, operations=None):
        super().__init__()
        self.norm_final = operations.LayerNorm(hidden_size, elementwise_affine=False, eps=1e-6, dtype=dtype, device=device)
-        self.linear = operations.Linear(hidden_size, patch_size * patch_size * out_channels, bias=bias, dtype=dtype, device=device)
-        self.adaLN_modulation = nn.Sequential(nn.SiLU(), operations.Linear(hidden_size, 2 * hidden_size, bias=bias, dtype=dtype, device=device))
+        self.linear = operations.Linear(hidden_size, patch_size * patch_size * out_channels, bias=True, dtype=dtype, device=device)
+        self.adaLN_modulation = nn.Sequential(nn.SiLU(), operations.Linear(hidden_size, 2 * hidden_size, bias=True, dtype=dtype, device=device))

    def forward(self, x: Tensor, vec: Tensor, modulation_dims=None) -> Tensor:
        if vec.ndim == 2:
--- a/comfy/ldm/flux/math.py
+++ b/comfy/ldm/flux/math.py
@@ -7,8 +7,15 @@ import comfy.model_management


 def attention(q: Tensor, k: Tensor, v: Tensor, pe: Tensor, mask=None, transformer_options={}) -> Tensor:
+    q_shape = q.shape
+    k_shape = k.shape
+
    if pe is not None:
-        q, k = apply_rope(q, k, pe)
+        q = q.to(dtype=pe.dtype).reshape(*q.shape[:-1], -1, 1, 2)
+        k = k.to(dtype=pe.dtype).reshape(*k.shape[:-1], -1, 1, 2)
+        q = (pe[..., 0] * q[..., 0] + pe[..., 1] * q[..., 1]).reshape(*q_shape).type_as(v)
+        k = (pe[..., 0] * k[..., 0] + pe[..., 1] * k[..., 1]).reshape(*k_shape).type_as(v)
+
    heads = q.shape[1]
    x = optimized_attention(q, k, v, heads, skip_reshape=True, mask=mask, transformer_options=transformer_options)
    return x
--- a/comfy/ldm/flux/model.py
+++ b/comfy/ldm/flux/model.py
@@ -7,6 +7,7 @@ from torch import Tensor, nn
 from einops import rearrange, repeat
 import comfy.ldm.common_dit
 import comfy.patcher_extension
+from comfy.ldm.flipflop_transformer import FlipFlopModule

 from .layers import (
    DoubleStreamBlock,
@@ -15,7 +16,6 @@ from .layers import (
    MLPEmbedder,
    SingleStreamBlock,
    timestep_embedding,
-    Modulation
 )

@dataclass
@@ -34,20 +34,15 @@ class FluxParams:
    patch_size: int
    qkv_bias: bool
    guidance_embed: bool
-    global_modulation: bool = False
-    mlp_silu_act: bool = False
-    ops_bias: bool = True
-    default_ref_method: str = "offset"
-    ref_index_scale: float = 1.0


-class Flux(nn.Module):
+class Flux(FlipFlopModule):
    """
    Transformer model for flow matching on sequences.
    """

    def __init__(self, image_model=None, final_layer=True, dtype=None, device=None, operations=None, **kwargs):
-        super().__init__()
+        super().__init__(("double_blocks", "single_blocks"))
        self.dtype = dtype
        params = FluxParams(**kwargs)
        self.params = params
@@ -64,17 +59,13 @@ class Flux(nn.Module):
        self.hidden_size = params.hidden_size
        self.num_heads = params.num_heads
        self.pe_embedder = EmbedND(dim=pe_dim, theta=params.theta, axes_dim=params.axes_dim)
-        self.img_in = operations.Linear(self.in_channels, self.hidden_size, bias=params.ops_bias, dtype=dtype, device=device)
-        self.time_in = MLPEmbedder(in_dim=256, hidden_dim=self.hidden_size, bias=params.ops_bias, dtype=dtype, device=device, operations=operations)
-        if params.vec_in_dim is not None:
-            self.vector_in = MLPEmbedder(params.vec_in_dim, self.hidden_size, dtype=dtype, device=device, operations=operations)
-        else:
-            self.vector_in = None
-
+        self.img_in = operations.Linear(self.in_channels, self.hidden_size, bias=True, dtype=dtype, device=device)
+        self.time_in = MLPEmbedder(in_dim=256, hidden_dim=self.hidden_size, dtype=dtype, device=device, operations=operations)
+        self.vector_in = MLPEmbedder(params.vec_in_dim, self.hidden_size, dtype=dtype, device=device, operations=operations)
        self.guidance_in = (
-            MLPEmbedder(in_dim=256, hidden_dim=self.hidden_size, bias=params.ops_bias, dtype=dtype, device=device, operations=operations) if params.guidance_embed else nn.Identity()
+            MLPEmbedder(in_dim=256, hidden_dim=self.hidden_size, dtype=dtype, device=device, operations=operations) if params.guidance_embed else nn.Identity()
        )
-        self.txt_in = operations.Linear(params.context_in_dim, self.hidden_size, bias=params.ops_bias, dtype=dtype, device=device)
+        self.txt_in = operations.Linear(params.context_in_dim, self.hidden_size, dtype=dtype, device=device)

        self.double_blocks = nn.ModuleList(
            [
@@ -83,9 +74,6 @@ class Flux(nn.Module):
                    self.num_heads,
                    mlp_ratio=params.mlp_ratio,
                    qkv_bias=params.qkv_bias,
-                    modulation=params.global_modulation is False,
-                    mlp_silu_act=params.mlp_silu_act,
-                    proj_bias=params.ops_bias,
                    dtype=dtype, device=device, operations=operations
                )
                for _ in range(params.depth)
@@ -94,30 +82,79 @@ class Flux(nn.Module):

        self.single_blocks = nn.ModuleList(
            [
-                SingleStreamBlock(self.hidden_size, self.num_heads, mlp_ratio=params.mlp_ratio, modulation=params.global_modulation is False, mlp_silu_act=params.mlp_silu_act, bias=params.ops_bias, dtype=dtype, device=device, operations=operations)
+                SingleStreamBlock(self.hidden_size, self.num_heads, mlp_ratio=params.mlp_ratio, dtype=dtype, device=device, operations=operations)
                for _ in range(params.depth_single_blocks)
            ]
        )

        if final_layer:
-            self.final_layer = LastLayer(self.hidden_size, 1, self.out_channels, bias=params.ops_bias, dtype=dtype, device=device, operations=operations)
+            self.final_layer = LastLayer(self.hidden_size, 1, self.out_channels, dtype=dtype, device=device, operations=operations)

-        if params.global_modulation:
-            self.double_stream_modulation_img = Modulation(
-                self.hidden_size,
-                double=True,
-                bias=False,
-                dtype=dtype, device=device, operations=operations
-            )
-            self.double_stream_modulation_txt = Modulation(
-                self.hidden_size,
-                double=True,
-                bias=False,
-                dtype=dtype, device=device, operations=operations
-            )
-            self.single_stream_modulation = Modulation(
-                self.hidden_size, double=False, bias=False, dtype=dtype, device=device, operations=operations
-            )
+    def indiv_double_block_fwd(self, i, block, img, txt, vec, pe, attn_mask, control, blocks_replace, transformer_options):
+        if ("double_block", i) in blocks_replace:
+            def block_wrap(args):
+                out = {}
+                out["img"], out["txt"] = block(img=args["img"],
+                                               txt=args["txt"],
+                                               vec=args["vec"],
+                                               pe=args["pe"],
+                                               attn_mask=args.get("attn_mask"),
+                                               transformer_options=args.get("transformer_options"))
+                return out
+
+            out = blocks_replace[("double_block", i)]({"img": img,
+                                                       "txt": txt,
+                                                       "vec": vec,
+                                                       "pe": pe,
+                                                       "attn_mask": attn_mask,
+                                                       "transformer_options": transformer_options},
+                                                       {"original_block": block_wrap})
+            txt = out["txt"]
+            img = out["img"]
+        else:
+            img, txt = block(img=img,
+                             txt=txt,
+                             vec=vec,
+                             pe=pe,
+                             attn_mask=attn_mask,
+                             transformer_options=transformer_options)
+
+        if control is not None: # Controlnet
+            control_i = control.get("input")
+            if i < len(control_i):
+                add = control_i[i]
+                if add is not None:
+                    img[:, :add.shape[1]] += add
+        return img, txt
+
+    def indiv_single_block_fwd(self, i, block, img, txt, vec, pe, attn_mask, control, blocks_replace, transformer_options):
+        if ("single_block", i) in blocks_replace:
+            def block_wrap(args):
+                out = {}
+                out["img"] = block(args["img"],
+                                   vec=args["vec"],
+                                   pe=args["pe"],
+                                   attn_mask=args.get("attn_mask"),
+                                   transformer_options=args.get("transformer_options"))
+                return out
+
+            out = blocks_replace[("single_block", i)]({"img": img,
+                                                       "vec": vec,
+                                                       "pe": pe,
+                                                       "attn_mask": attn_mask,
+                                                       "transformer_options": transformer_options},
+                                                       {"original_block": block_wrap})
+            img = out["img"]
+        else:
+            img = block(img, vec=vec, pe=pe, attn_mask=attn_mask, transformer_options=transformer_options)
+
+        if control is not None: # Controlnet
+            control_o = control.get("output")
+            if i < len(control_o):
+                add = control_o[i]
+                if add is not None:
+                    img[:, txt.shape[1] : txt.shape[1] + add.shape[1], ...] += add
+        return img

    def forward_orig(
        self,
@@ -133,6 +170,9 @@ class Flux(nn.Module):
        attn_mask: Tensor = None,
    ) -> Tensor:

+        if y is None:
+            y = torch.zeros((img.shape[0], self.params.vec_in_dim), device=img.device, dtype=img.dtype)
+
        patches = transformer_options.get("patches", {})
        patches_replace = transformer_options.get("patches_replace", {})
        if img.ndim != 3 or txt.ndim != 3:
@@ -145,17 +185,9 @@ class Flux(nn.Module):
            if guidance is not None:
                vec = vec + self.guidance_in(timestep_embedding(guidance, 256).to(img.dtype))

-        if self.vector_in is not None:
-            if y is None:
-                y = torch.zeros((img.shape[0], self.params.vec_in_dim), device=img.device, dtype=img.dtype)
-            vec = vec + self.vector_in(y[:, :self.params.vec_in_dim])
-
+        vec = vec + self.vector_in(y[:, :self.params.vec_in_dim])
        txt = self.txt_in(txt)

-        vec_orig = vec
-        if self.params.global_modulation:
-            vec = (self.double_stream_modulation_img(vec_orig), self.double_stream_modulation_txt(vec_orig))
-
        if "post_input" in patches:
            for p in patches["post_input"]:
                out = p({"img": img, "txt": txt, "img_ids": img_ids, "txt_ids": txt_ids})
@@ -171,90 +203,23 @@ class Flux(nn.Module):
            pe = None

        blocks_replace = patches_replace.get("dit", {})
-        transformer_options["total_blocks"] = len(self.double_blocks)
-        transformer_options["block_type"] = "double"
-        for i, block in enumerate(self.double_blocks):
-            transformer_options["block_index"] = i
-            if ("double_block", i) in blocks_replace:
-                def block_wrap(args):
-                    out = {}
-                    out["img"], out["txt"] = block(img=args["img"],
-                                                   txt=args["txt"],
-                                                   vec=args["vec"],
-                                                   pe=args["pe"],
-                                                   attn_mask=args.get("attn_mask"),
-                                                   transformer_options=args.get("transformer_options"))
-                    return out
-
-                out = blocks_replace[("double_block", i)]({"img": img,
-                                                           "txt": txt,
-                                                           "vec": vec,
-                                                           "pe": pe,
-                                                           "attn_mask": attn_mask,
-                                                           "transformer_options": transformer_options},
-                                                          {"original_block": block_wrap})
-                txt = out["txt"]
-                img = out["img"]
-            else:
-                img, txt = block(img=img,
-                                 txt=txt,
-                                 vec=vec,
-                                 pe=pe,
-                                 attn_mask=attn_mask,
-                                 transformer_options=transformer_options)
-
-            if control is not None: # Controlnet
-                control_i = control.get("input")
-                if i < len(control_i):
-                    add = control_i[i]
-                    if add is not None:
-                        img[:, :add.shape[1]] += add
+        # execute double blocks
+        img, txt = self.execute_blocks("double_blocks", self.indiv_double_block_fwd, (img, txt), vec, pe, attn_mask, control, blocks_replace, transformer_options)

        if img.dtype == torch.float16:
            img = torch.nan_to_num(img, nan=0.0, posinf=65504, neginf=-65504)

        img = torch.cat((txt, img), 1)

-        if self.params.global_modulation:
-            vec, _ = self.single_stream_modulation(vec_orig)
-
-        transformer_options["total_blocks"] = len(self.single_blocks)
-        transformer_options["block_type"] = "single"
-        for i, block in enumerate(self.single_blocks):
-            transformer_options["block_index"] = i
-            if ("single_block", i) in blocks_replace:
-                def block_wrap(args):
-                    out = {}
-                    out["img"] = block(args["img"],
-                                       vec=args["vec"],
-                                       pe=args["pe"],
-                                       attn_mask=args.get("attn_mask"),
-                                       transformer_options=args.get("transformer_options"))
-                    return out
-
-                out = blocks_replace[("single_block", i)]({"img": img,
-                                                           "vec": vec,
-                                                           "pe": pe,
-                                                           "attn_mask": attn_mask,
-                                                           "transformer_options": transformer_options},
-                                                          {"original_block": block_wrap})
-                img = out["img"]
-            else:
-                img = block(img, vec=vec, pe=pe, attn_mask=attn_mask, transformer_options=transformer_options)
-
-            if control is not None: # Controlnet
-                control_o = control.get("output")
-                if i < len(control_o):
-                    add = control_o[i]
-                    if add is not None:
-                        img[:, txt.shape[1] : txt.shape[1] + add.shape[1], ...] += add
+        # execute single blocks
+        img = self.execute_blocks("single_blocks", self.indiv_single_block_fwd, img, txt, vec, pe, attn_mask, control, blocks_replace, transformer_options)

        img = img[:, txt.shape[1] :, ...]

-        img = self.final_layer(img, vec_orig)  # (N, T, patch_size ** 2 * out_channels)
+        img = self.final_layer(img, vec)  # (N, T, patch_size ** 2 * out_channels)
        return img

-    def process_img(self, x, index=0, h_offset=0, w_offset=0, transformer_options={}):
+    def process_img(self, x, index=0, h_offset=0, w_offset=0):
        bs, c, h, w = x.shape
        patch_size = self.patch_size
        x = comfy.ldm.common_dit.pad_to_patch_size(x, (patch_size, patch_size))
@@ -266,22 +231,10 @@ class Flux(nn.Module):
        h_offset = ((h_offset + (patch_size // 2)) // patch_size)
        w_offset = ((w_offset + (patch_size // 2)) // patch_size)

-        steps_h = h_len
-        steps_w = w_len
-
-        rope_options = transformer_options.get("rope_options", None)
-        if rope_options is not None:
-            h_len = (h_len - 1.0) * rope_options.get("scale_y", 1.0) + 1.0
-            w_len = (w_len - 1.0) * rope_options.get("scale_x", 1.0) + 1.0
-
-            index += rope_options.get("shift_t", 0.0)
-            h_offset += rope_options.get("shift_y", 0.0)
-            w_offset += rope_options.get("shift_x", 0.0)
-
-        img_ids = torch.zeros((steps_h, steps_w, len(self.params.axes_dim)), device=x.device, dtype=torch.float32)
+        img_ids = torch.zeros((h_len, w_len, 3), device=x.device, dtype=x.dtype)
        img_ids[:, :, 0] = img_ids[:, :, 1] + index
-        img_ids[:, :, 1] = img_ids[:, :, 1] + torch.linspace(h_offset, h_len - 1 + h_offset, steps=steps_h, device=x.device, dtype=torch.float32).unsqueeze(1)
-        img_ids[:, :, 2] = img_ids[:, :, 2] + torch.linspace(w_offset, w_len - 1 + w_offset, steps=steps_w, device=x.device, dtype=torch.float32).unsqueeze(0)
+        img_ids[:, :, 1] = img_ids[:, :, 1] + torch.linspace(h_offset, h_len - 1 + h_offset, steps=h_len, device=x.device, dtype=x.dtype).unsqueeze(1)
+        img_ids[:, :, 2] = img_ids[:, :, 2] + torch.linspace(w_offset, w_len - 1 + w_offset, steps=w_len, device=x.device, dtype=x.dtype).unsqueeze(0)
        return img, repeat(img_ids, "h w c -> b (h w) c", b=bs)

    def forward(self, x, timestep, context, y=None, guidance=None, ref_latents=None, control=None, transformer_options={}, **kwargs):
@@ -297,16 +250,16 @@ class Flux(nn.Module):

        h_len = ((h_orig + (patch_size // 2)) // patch_size)
        w_len = ((w_orig + (patch_size // 2)) // patch_size)
-        img, img_ids = self.process_img(x, transformer_options=transformer_options)
+        img, img_ids = self.process_img(x)
        img_tokens = img.shape[1]
        if ref_latents is not None:
            h = 0
            w = 0
            index = 0
-            ref_latents_method = kwargs.get("ref_latents_method", self.params.default_ref_method)
+            ref_latents_method = kwargs.get("ref_latents_method", "offset")
            for ref in ref_latents:
                if ref_latents_method == "index":
-                    index += self.params.ref_index_scale
+                    index += 1
                    h_offset = 0
                    w_offset = 0
                elif ref_latents_method == "uxo":
@@ -330,11 +283,7 @@ class Flux(nn.Module):
                img = torch.cat([img, kontext], dim=1)
                img_ids = torch.cat([img_ids, kontext_ids], dim=1)

-        txt_ids = torch.zeros((bs, context.shape[1], len(self.params.axes_dim)), device=x.device, dtype=torch.float32)
-
-        if len(self.params.axes_dim) == 4: # Flux 2
-            txt_ids[:, :, 3] = torch.linspace(0, context.shape[1] - 1, steps=context.shape[1], device=x.device, dtype=torch.float32)
-
+        txt_ids = torch.zeros((bs, context.shape[1], 3), device=x.device, dtype=x.dtype)
        out = self.forward_orig(img, img_ids, context, txt_ids, timestep, y, guidance, control, transformer_options, attn_mask=kwargs.get("attention_mask", None))
        out = out[:, :img_tokens]
-        return rearrange(out, "b (h w) (c ph pw) -> b c (h ph) (w pw)", h=h_len, w=w_len, ph=self.patch_size, pw=self.patch_size)[:,:,:h_orig,:w_orig]
+        return rearrange(out, "b (h w) (c ph pw) -> b c (h ph) (w pw)", h=h_len, w=w_len, ph=2, pw=2)[:,:,:h_orig,:w_orig]
--- a/comfy/ldm/hunyuan_video/model.py
+++ b/comfy/ldm/hunyuan_video/model.py
@@ -6,6 +6,7 @@ import comfy.ldm.flux.layers
 import comfy.ldm.modules.diffusionmodules.mmdit
 from comfy.ldm.modules.attention import optimized_attention

+
 from dataclasses import dataclass
 from einops import repeat

@@ -41,8 +42,6 @@ class HunyuanVideoParams:
    guidance_embed: bool
    byt5: bool
    meanflow: bool
-    use_cond_type_embedding: bool
-    vision_in_dim: int


 class SelfAttentionRef(nn.Module):
@@ -158,10 +157,7 @@ class TokenRefiner(nn.Module):
        t = self.t_embedder(timestep_embedding(timesteps, 256, time_factor=1.0).to(x.dtype))
        # m = mask.float().unsqueeze(-1)
        # c = (x.float() * m).sum(dim=1) / m.sum(dim=1) #TODO: the following works when the x.shape is the same length as the tokens but might break otherwise
-        if x.dtype == torch.float16:
-            c = x.float().sum(dim=1) / x.shape[1]
-        else:
-            c = x.sum(dim=1) / x.shape[1]
+        c = x.sum(dim=1) / x.shape[1]

        c = t + self.c_embedder(c.to(x.dtype))
        x = self.input_embedder(x)
@@ -200,15 +196,11 @@ class HunyuanVideo(nn.Module):
    def __init__(self, image_model=None, final_layer=True, dtype=None, device=None, operations=None, **kwargs):
        super().__init__()
        self.dtype = dtype
-        operation_settings = {"operations": operations, "device": device, "dtype": dtype}
-
        params = HunyuanVideoParams(**kwargs)
        self.params = params
        self.patch_size = params.patch_size
        self.in_channels = params.in_channels
        self.out_channels = params.out_channels
-        self.use_cond_type_embedding = params.use_cond_type_embedding
-        self.vision_in_dim = params.vision_in_dim
        if params.hidden_size % params.num_heads != 0:
            raise ValueError(
                f"Hidden size {params.hidden_size} must be divisible by num_heads {params.num_heads}"
@@ -274,18 +266,6 @@ class HunyuanVideo(nn.Module):
        if final_layer:
            self.final_layer = LastLayer(self.hidden_size, self.patch_size[-1], self.out_channels, dtype=dtype, device=device, operations=operations)

-        # HunyuanVideo 1.5 specific modules
-        if self.vision_in_dim is not None:
-            from comfy.ldm.wan.model import MLPProj
-            self.vision_in = MLPProj(in_dim=self.vision_in_dim, out_dim=self.hidden_size, operation_settings=operation_settings)
-        else:
-            self.vision_in = None
-        if self.use_cond_type_embedding:
-            # 0: text_encoder feature 1: byt5 feature 2: vision_encoder feature
-            self.cond_type_embedding = nn.Embedding(3, self.hidden_size)
-        else:
-            self.cond_type_embedding = None
-
    def forward_orig(
        self,
        img: Tensor,
@@ -296,7 +276,6 @@ class HunyuanVideo(nn.Module):
        timesteps: Tensor,
        y: Tensor = None,
        txt_byt5=None,
-        clip_fea=None,
        guidance: Tensor = None,
        guiding_frame_index=None,
        ref_latent=None,
@@ -352,31 +331,12 @@ class HunyuanVideo(nn.Module):

        txt = self.txt_in(txt, timesteps, txt_mask, transformer_options=transformer_options)

-        if self.cond_type_embedding is not None:
-            self.cond_type_embedding.to(txt.device)
-            cond_emb = self.cond_type_embedding(torch.zeros_like(txt[:, :, 0], device=txt.device, dtype=torch.long))
-            txt = txt + cond_emb.to(txt.dtype)
-
        if self.byt5_in is not None and txt_byt5 is not None:
            txt_byt5 = self.byt5_in(txt_byt5)
-            if self.cond_type_embedding is not None:
-                cond_emb = self.cond_type_embedding(torch.ones_like(txt_byt5[:, :, 0], device=txt_byt5.device, dtype=torch.long))
-                txt_byt5 = txt_byt5 + cond_emb.to(txt_byt5.dtype)
-                txt = torch.cat((txt_byt5, txt), dim=1) # byt5 first for HunyuanVideo1.5
-            else:
-                txt = torch.cat((txt, txt_byt5), dim=1)
            txt_byt5_ids = torch.zeros((txt_ids.shape[0], txt_byt5.shape[1], txt_ids.shape[-1]), device=txt_ids.device, dtype=txt_ids.dtype)
+            txt = torch.cat((txt, txt_byt5), dim=1)
            txt_ids = torch.cat((txt_ids, txt_byt5_ids), dim=1)

-        if clip_fea is not None:
-            txt_vision_states = self.vision_in(clip_fea)
-            if self.cond_type_embedding is not None:
-                cond_emb = self.cond_type_embedding(2 * torch.ones_like(txt_vision_states[:, :, 0], dtype=torch.long, device=txt_vision_states.device))
-                txt_vision_states = txt_vision_states + cond_emb
-            txt = torch.cat((txt_vision_states.to(txt.dtype), txt), dim=1)
-            extra_txt_ids = torch.zeros((txt_ids.shape[0], txt_vision_states.shape[1], txt_ids.shape[-1]), device=txt_ids.device, dtype=txt_ids.dtype)
-            txt_ids = torch.cat((txt_ids, extra_txt_ids), dim=1)
-
        ids = torch.cat((img_ids, txt_ids), dim=1)
        pe = self.pe_embedder(ids)

@@ -389,10 +349,7 @@ class HunyuanVideo(nn.Module):
            attn_mask = None

        blocks_replace = patches_replace.get("dit", {})
-        transformer_options["total_blocks"] = len(self.double_blocks)
-        transformer_options["block_type"] = "double"
        for i, block in enumerate(self.double_blocks):
-            transformer_options["block_index"] = i
            if ("double_block", i) in blocks_replace:
                def block_wrap(args):
                    out = {}
@@ -414,10 +371,7 @@ class HunyuanVideo(nn.Module):

        img = torch.cat((img, txt), 1)

-        transformer_options["total_blocks"] = len(self.single_blocks)
-        transformer_options["block_type"] = "single"
        for i, block in enumerate(self.single_blocks):
-            transformer_options["block_index"] = i
            if ("single_block", i) in blocks_replace:
                def block_wrap(args):
                    out = {}
@@ -476,14 +430,14 @@ class HunyuanVideo(nn.Module):
        img_ids[:, :, 1] = img_ids[:, :, 1] + torch.linspace(0, w_len - 1, steps=w_len, device=x.device, dtype=x.dtype).unsqueeze(0)
        return repeat(img_ids, "h w c -> b (h w) c", b=bs)

-    def forward(self, x, timestep, context, y=None, txt_byt5=None, clip_fea=None, guidance=None, attention_mask=None, guiding_frame_index=None, ref_latent=None, disable_time_r=False, control=None, transformer_options={}, **kwargs):
+    def forward(self, x, timestep, context, y=None, txt_byt5=None, guidance=None, attention_mask=None, guiding_frame_index=None, ref_latent=None, disable_time_r=False, control=None, transformer_options={}, **kwargs):
        return comfy.patcher_extension.WrapperExecutor.new_class_executor(
            self._forward,
            self,
            comfy.patcher_extension.get_all_wrappers(comfy.patcher_extension.WrappersMP.DIFFUSION_MODEL, transformer_options)
-        ).execute(x, timestep, context, y, txt_byt5, clip_fea, guidance, attention_mask, guiding_frame_index, ref_latent, disable_time_r, control, transformer_options, **kwargs)
+        ).execute(x, timestep, context, y, txt_byt5, guidance, attention_mask, guiding_frame_index, ref_latent, disable_time_r, control, transformer_options, **kwargs)

-    def _forward(self, x, timestep, context, y=None, txt_byt5=None, clip_fea=None, guidance=None, attention_mask=None, guiding_frame_index=None, ref_latent=None, disable_time_r=False, control=None, transformer_options={}, **kwargs):
+    def _forward(self, x, timestep, context, y=None, txt_byt5=None, guidance=None, attention_mask=None, guiding_frame_index=None, ref_latent=None, disable_time_r=False, control=None, transformer_options={}, **kwargs):
        bs = x.shape[0]
        if len(self.patch_size) == 3:
            img_ids = self.img_ids(x)
@@ -491,5 +445,5 @@ class HunyuanVideo(nn.Module):
        else:
            img_ids = self.img_ids_2d(x)
            txt_ids = torch.zeros((bs, context.shape[1], 2), device=x.device, dtype=x.dtype)
-        out = self.forward_orig(x, img_ids, context, txt_ids, attention_mask, timestep, y, txt_byt5, clip_fea, guidance, guiding_frame_index, ref_latent, disable_time_r=disable_time_r, control=control, transformer_options=transformer_options)
+        out = self.forward_orig(x, img_ids, context, txt_ids, attention_mask, timestep, y, txt_byt5, guidance, guiding_frame_index, ref_latent, disable_time_r=disable_time_r, control=control, transformer_options=transformer_options)
        return out
--- a/comfy/ldm/hunyuan_video/upsampler.py
+++ b/comfy/ldm/hunyuan_video/upsampler.py
@@ -1,120 +0,0 @@
-import torch
-import torch.nn as nn
-import torch.nn.functional as F
-from comfy.ldm.hunyuan_video.vae_refiner import RMS_norm, ResnetBlock, VideoConv3d
-import model_management, model_patcher
-
-class SRResidualCausalBlock3D(nn.Module):
-    def __init__(self, channels: int):
-        super().__init__()
-        self.block = nn.Sequential(
-            VideoConv3d(channels, channels, kernel_size=3),
-            nn.SiLU(inplace=True),
-            VideoConv3d(channels, channels, kernel_size=3),
-            nn.SiLU(inplace=True),
-            VideoConv3d(channels, channels, kernel_size=3),
-        )
-
-    def forward(self, x: torch.Tensor) -> torch.Tensor:
-        return x + self.block(x)
-
-class SRModel3DV2(nn.Module):
-    def __init__(
-        self,
-        in_channels: int,
-        out_channels: int,
-        hidden_channels: int = 64,
-        num_blocks: int = 6,
-        global_residual: bool = False,
-    ):
-        super().__init__()
-        self.in_conv = VideoConv3d(in_channels, hidden_channels, kernel_size=3)
-        self.blocks = nn.ModuleList([SRResidualCausalBlock3D(hidden_channels) for _ in range(num_blocks)])
-        self.out_conv = VideoConv3d(hidden_channels, out_channels, kernel_size=3)
-        self.global_residual = bool(global_residual)
-
-    def forward(self, x: torch.Tensor) -> torch.Tensor:
-        residual = x
-        y = self.in_conv(x)
-        for blk in self.blocks:
-            y = blk(y)
-        y = self.out_conv(y)
-        if self.global_residual and (y.shape == residual.shape):
-            y = y + residual
-        return y
-
-
-class Upsampler(nn.Module):
-    def __init__(
-        self,
-        z_channels: int,
-        out_channels: int,
-        block_out_channels: tuple[int, ...],
-        num_res_blocks: int = 2,
-    ):
-        super().__init__()
-        self.num_res_blocks = num_res_blocks
-        self.block_out_channels = block_out_channels
-        self.z_channels = z_channels
-
-        ch = block_out_channels[0]
-        self.conv_in = VideoConv3d(z_channels, ch, kernel_size=3)
-
-        self.up = nn.ModuleList()
-
-        for i, tgt in enumerate(block_out_channels):
-            stage = nn.Module()
-            stage.block = nn.ModuleList([ResnetBlock(in_channels=ch if j == 0 else tgt,
-                                                    out_channels=tgt,
-                                                    temb_channels=0,
-                                                    conv_shortcut=False,
-                                                    conv_op=VideoConv3d, norm_op=RMS_norm)
-                                        for j in range(num_res_blocks + 1)])
-            ch = tgt
-            self.up.append(stage)
-
-        self.norm_out = RMS_norm(ch)
-        self.conv_out = VideoConv3d(ch, out_channels, kernel_size=3)
-
-    def forward(self, z):
-        """
-        Args:
-            z: (B, C, T, H, W)
-            target_shape: (H, W)
-        """
-        # z to block_in
-        repeats = self.block_out_channels[0] // (self.z_channels)
-        x = self.conv_in(z) + z.repeat_interleave(repeats=repeats, dim=1)
-
-        # upsampling
-        for stage in self.up:
-            for blk in stage.block:
-                x = blk(x)
-
-        out = self.conv_out(F.silu(self.norm_out(x)))
-        return out
-
-UPSAMPLERS = {
-    "720p": SRModel3DV2,
-    "1080p": Upsampler,
-}
-
-class HunyuanVideo15SRModel():
-    def __init__(self, model_type, config):
-        self.load_device = model_management.vae_device()
-        offload_device = model_management.vae_offload_device()
-        self.dtype = model_management.vae_dtype(self.load_device)
-        self.model_class = UPSAMPLERS.get(model_type)
-        self.model = self.model_class(**config).eval()
-
-        self.patcher = model_patcher.ModelPatcher(self.model, load_device=self.load_device, offload_device=offload_device)
-
-    def load_sd(self, sd):
-        return self.model.load_state_dict(sd, strict=True)
-
-    def get_sd(self):
-        return self.model.state_dict()
-
-    def resample_latent(self, latent):
-        model_management.load_model_gpu(self.patcher)
-        return self.model(latent.to(self.load_device))
--- a/comfy/ldm/hunyuan_video/vae_refiner.py
+++ b/comfy/ldm/hunyuan_video/vae_refiner.py
@@ -4,40 +4,8 @@ import torch.nn.functional as F
 from comfy.ldm.modules.diffusionmodules.model import ResnetBlock, AttnBlock, VideoConv3d, Normalize
 import comfy.ops
 import comfy.ldm.models.autoencoder
-import comfy.model_management
 ops = comfy.ops.disable_weight_init

-class NoPadConv3d(nn.Module):
-    def __init__(self, n_channels, out_channels, kernel_size, stride=1, dilation=1, padding=0, **kwargs):
-        super().__init__()
-        self.conv = ops.Conv3d(n_channels, out_channels, kernel_size, stride=stride, dilation=dilation, **kwargs)
-
-    def forward(self, x):
-        return self.conv(x)
-
-
-def conv_carry_causal_3d(xl, op, conv_carry_in=None, conv_carry_out=None):
-
-    x = xl[0]
-    xl.clear()
-
-    if conv_carry_out is not None:
-        to_push = x[:, :, -2:, :, :].clone()
-        conv_carry_out.append(to_push)
-
-    if isinstance(op, NoPadConv3d):
-        if conv_carry_in is None:
-            x = torch.nn.functional.pad(x, (1, 1, 1, 1, 2, 0), mode = 'replicate')
-        else:
-            carry_len = conv_carry_in[0].shape[2]
-            x = torch.cat([conv_carry_in.pop(0), x], dim=2)
-            x = torch.nn.functional.pad(x, (1, 1, 1, 1, 2 - carry_len, 0), mode = 'replicate')
-
-    out = op(x)
-
-    return out
-
-
 class RMS_norm(nn.Module):
    def __init__(self, dim):
        super().__init__()
@@ -46,7 +14,7 @@ class RMS_norm(nn.Module):
        self.gamma = nn.Parameter(torch.empty(shape))

    def forward(self, x):
-        return F.normalize(x, dim=1) * self.scale * comfy.model_management.cast_to(self.gamma, dtype=x.dtype, device=x.device)
+        return F.normalize(x, dim=1) * self.scale * self.gamma

 class DnSmpl(nn.Module):
    def __init__(self, ic, oc, tds=True, refiner_vae=True, op=VideoConv3d):
@@ -59,12 +27,11 @@ class DnSmpl(nn.Module):
        self.tds = tds
        self.gs = fct * ic // oc

-    def forward(self, x, conv_carry_in=None, conv_carry_out=None):
+    def forward(self, x):
        r1 = 2 if self.tds else 1
-        h = conv_carry_causal_3d([x], self.conv, conv_carry_in, conv_carry_out)
-
-        if self.tds and self.refiner_vae and conv_carry_in is None:
+        h = self.conv(x)

+        if self.tds and self.refiner_vae:
            hf = h[:, :, :1, :, :]
            b, c, f, ht, wd = hf.shape
            hf = hf.reshape(b, c, f, ht // 2, 2, wd // 2, 2)
@@ -72,7 +39,14 @@ class DnSmpl(nn.Module):
            hf = hf.reshape(b, 2 * 2 * c, f, ht // 2, wd // 2)
            hf = torch.cat([hf, hf], dim=1)

-            h = h[:, :, 1:, :, :]
+            hn = h[:, :, 1:, :, :]
+            b, c, frms, ht, wd = hn.shape
+            nf = frms // r1
+            hn = hn.reshape(b, c, nf, r1, ht // 2, 2, wd // 2, 2)
+            hn = hn.permute(0, 3, 5, 7, 1, 2, 4, 6)
+            hn = hn.reshape(b, r1 * 2 * 2 * c, nf, ht // 2, wd // 2)
+
+            h = torch.cat([hf, hn], dim=2)

            xf = x[:, :, :1, :, :]
            b, ci, f, ht, wd = xf.shape
@@ -80,32 +54,34 @@ class DnSmpl(nn.Module):
            xf = xf.permute(0, 4, 6, 1, 2, 3, 5)
            xf = xf.reshape(b, 2 * 2 * ci, f, ht // 2, wd // 2)
            B, C, T, H, W = xf.shape
-            xf = xf.view(B, hf.shape[1], self.gs // 2, T, H, W).mean(dim=2)
+            xf = xf.view(B, h.shape[1], self.gs // 2, T, H, W).mean(dim=2)

-            x = x[:, :, 1:, :, :]
+            xn = x[:, :, 1:, :, :]
+            b, ci, frms, ht, wd = xn.shape
+            nf = frms // r1
+            xn = xn.reshape(b, ci, nf, r1, ht // 2, 2, wd // 2, 2)
+            xn = xn.permute(0, 3, 5, 7, 1, 2, 4, 6)
+            xn = xn.reshape(b, r1 * 2 * 2 * ci, nf, ht // 2, wd // 2)
+            B, C, T, H, W = xn.shape
+            xn = xn.view(B, h.shape[1], self.gs, T, H, W).mean(dim=2)
+            sc = torch.cat([xf, xn], dim=2)
+        else:
+            b, c, frms, ht, wd = h.shape

-        if h.shape[2] == 0:
-            return hf + xf
+            nf = frms // r1
+            h = h.reshape(b, c, nf, r1, ht // 2, 2, wd // 2, 2)
+            h = h.permute(0, 3, 5, 7, 1, 2, 4, 6)
+            h = h.reshape(b, r1 * 2 * 2 * c, nf, ht // 2, wd // 2)

-        b, c, frms, ht, wd = h.shape
-        nf = frms // r1
-        h = h.reshape(b, c, nf, r1, ht // 2, 2, wd // 2, 2)
-        h = h.permute(0, 3, 5, 7, 1, 2, 4, 6)
-        h = h.reshape(b, r1 * 2 * 2 * c, nf, ht // 2, wd // 2)
+            b, ci, frms, ht, wd = x.shape
+            nf = frms // r1
+            sc = x.reshape(b, ci, nf, r1, ht // 2, 2, wd // 2, 2)
+            sc = sc.permute(0, 3, 5, 7, 1, 2, 4, 6)
+            sc = sc.reshape(b, r1 * 2 * 2 * ci, nf, ht // 2, wd // 2)
+            B, C, T, H, W = sc.shape
+            sc = sc.view(B, h.shape[1], self.gs, T, H, W).mean(dim=2)

-        b, ci, frms, ht, wd = x.shape
-        nf = frms // r1
-        x = x.reshape(b, ci, nf, r1, ht // 2, 2, wd // 2, 2)
-        x = x.permute(0, 3, 5, 7, 1, 2, 4, 6)
-        x = x.reshape(b, r1 * 2 * 2 * ci, nf, ht // 2, wd // 2)
-        B, C, T, H, W = x.shape
-        x = x.view(B, h.shape[1], self.gs, T, H, W).mean(dim=2)
-
-        if self.tds and self.refiner_vae and conv_carry_in is None:
-            h = torch.cat([hf, h], dim=2)
-            x = torch.cat([xf, x], dim=2)
-
-        return h + x
+        return h + sc


 class UpSmpl(nn.Module):
@@ -118,11 +94,11 @@ class UpSmpl(nn.Module):
        self.tus = tus
        self.rp = fct * oc // ic

-    def forward(self, x, conv_carry_in=None, conv_carry_out=None):
+    def forward(self, x):
        r1 = 2 if self.tus else 1
-        h = conv_carry_causal_3d([x], self.conv, conv_carry_in, conv_carry_out)
+        h = self.conv(x)

-        if self.tus and self.refiner_vae and conv_carry_in is None:
+        if self.tus and self.refiner_vae:
            hf = h[:, :, :1, :, :]
            b, c, f, ht, wd = hf.shape
            nc = c // (2 * 2)
@@ -131,7 +107,14 @@ class UpSmpl(nn.Module):
            hf = hf.reshape(b, nc, f, ht * 2, wd * 2)
            hf = hf[:, : hf.shape[1] // 2]

-            h = h[:, :, 1:, :, :]
+            hn = h[:, :, 1:, :, :]
+            b, c, frms, ht, wd = hn.shape
+            nc = c // (r1 * 2 * 2)
+            hn = hn.reshape(b, r1, 2, 2, nc, frms, ht, wd)
+            hn = hn.permute(0, 4, 5, 1, 6, 2, 7, 3)
+            hn = hn.reshape(b, nc, frms * r1, ht * 2, wd * 2)
+
+            h = torch.cat([hf, hn], dim=2)

            xf = x[:, :, :1, :, :]
            b, ci, f, ht, wd = xf.shape
@@ -142,43 +125,29 @@ class UpSmpl(nn.Module):
            xf = xf.permute(0, 3, 4, 5, 1, 6, 2)
            xf = xf.reshape(b, nc, f, ht * 2, wd * 2)

-            x = x[:, :, 1:, :, :]
+            xn = x[:, :, 1:, :, :]
+            xn = xn.repeat_interleave(repeats=self.rp, dim=1)
+            b, c, frms, ht, wd = xn.shape
+            nc = c // (r1 * 2 * 2)
+            xn = xn.reshape(b, r1, 2, 2, nc, frms, ht, wd)
+            xn = xn.permute(0, 4, 5, 1, 6, 2, 7, 3)
+            xn = xn.reshape(b, nc, frms * r1, ht * 2, wd * 2)
+            sc = torch.cat([xf, xn], dim=2)
+        else:
+            b, c, frms, ht, wd = h.shape
+            nc = c // (r1 * 2 * 2)
+            h = h.reshape(b, r1, 2, 2, nc, frms, ht, wd)
+            h = h.permute(0, 4, 5, 1, 6, 2, 7, 3)
+            h = h.reshape(b, nc, frms * r1, ht * 2, wd * 2)

-        b, c, frms, ht, wd = h.shape
-        nc = c // (r1 * 2 * 2)
-        h = h.reshape(b, r1, 2, 2, nc, frms, ht, wd)
-        h = h.permute(0, 4, 5, 1, 6, 2, 7, 3)
-        h = h.reshape(b, nc, frms * r1, ht * 2, wd * 2)
+            sc = x.repeat_interleave(repeats=self.rp, dim=1)
+            b, c, frms, ht, wd = sc.shape
+            nc = c // (r1 * 2 * 2)
+            sc = sc.reshape(b, r1, 2, 2, nc, frms, ht, wd)
+            sc = sc.permute(0, 4, 5, 1, 6, 2, 7, 3)
+            sc = sc.reshape(b, nc, frms * r1, ht * 2, wd * 2)

-        x = x.repeat_interleave(repeats=self.rp, dim=1)
-        b, c, frms, ht, wd = x.shape
-        nc = c // (r1 * 2 * 2)
-        x = x.reshape(b, r1, 2, 2, nc, frms, ht, wd)
-        x = x.permute(0, 4, 5, 1, 6, 2, 7, 3)
-        x = x.reshape(b, nc, frms * r1, ht * 2, wd * 2)
-
-        if self.tus and self.refiner_vae and conv_carry_in is None:
-            h = torch.cat([hf, h], dim=2)
-            x = torch.cat([xf, x], dim=2)
-
-        return h + x
-
-class HunyuanRefinerResnetBlock(ResnetBlock):
-    def __init__(self, in_channels, out_channels, conv_op=NoPadConv3d, norm_op=RMS_norm):
-        super().__init__(in_channels=in_channels, out_channels=out_channels, temb_channels=0, conv_op=conv_op, norm_op=norm_op)
-
-    def forward(self, x, conv_carry_in=None, conv_carry_out=None):
-        h = x
-        h = [ self.swish(self.norm1(x)) ]
-        h = conv_carry_causal_3d(h, self.conv1, conv_carry_in=conv_carry_in, conv_carry_out=conv_carry_out)
-
-        h = [ self.dropout(self.swish(self.norm2(h))) ]
-        h = conv_carry_causal_3d(h, self.conv2, conv_carry_in=conv_carry_in, conv_carry_out=conv_carry_out)
-
-        if self.in_channels != self.out_channels:
-            x = self.nin_shortcut(x)
-
-        return x+h
+        return h + sc

 class Encoder(nn.Module):
    def __init__(self, in_channels, z_channels, block_out_channels, num_res_blocks,
@@ -191,7 +160,7 @@ class Encoder(nn.Module):

        self.refiner_vae = refiner_vae
        if self.refiner_vae:
-            conv_op = NoPadConv3d
+            conv_op = VideoConv3d
            norm_op = RMS_norm
        else:
            conv_op = ops.Conv3d
@@ -206,9 +175,10 @@ class Encoder(nn.Module):

        for i, tgt in enumerate(block_out_channels):
            stage = nn.Module()
-            stage.block = nn.ModuleList([HunyuanRefinerResnetBlock(in_channels=ch if j == 0 else tgt,
-                                                                   out_channels=tgt,
-                                                                   conv_op=conv_op, norm_op=norm_op)
+            stage.block = nn.ModuleList([ResnetBlock(in_channels=ch if j == 0 else tgt,
+                                                     out_channels=tgt,
+                                                     temb_channels=0,
+                                                     conv_op=conv_op, norm_op=norm_op)
                                        for j in range(num_res_blocks)])
            ch = tgt
            if i < depth:
@@ -218,9 +188,9 @@ class Encoder(nn.Module):
            self.down.append(stage)

        self.mid = nn.Module()
-        self.mid.block_1 = HunyuanRefinerResnetBlock(in_channels=ch, out_channels=ch, conv_op=conv_op, norm_op=norm_op)
+        self.mid.block_1 = ResnetBlock(in_channels=ch, out_channels=ch, temb_channels=0, conv_op=conv_op, norm_op=norm_op)
        self.mid.attn_1 = AttnBlock(ch, conv_op=ops.Conv3d, norm_op=norm_op)
-        self.mid.block_2 = HunyuanRefinerResnetBlock(in_channels=ch, out_channels=ch, conv_op=conv_op, norm_op=norm_op)
+        self.mid.block_2 = ResnetBlock(in_channels=ch, out_channels=ch, temb_channels=0, conv_op=conv_op, norm_op=norm_op)

        self.norm_out = norm_op(ch)
        self.conv_out = conv_op(ch, z_channels << 1, 3, 1, 1)
@@ -231,50 +201,31 @@ class Encoder(nn.Module):
        if not self.refiner_vae and x.shape[2] == 1:
            x = x.expand(-1, -1, self.ffactor_temporal, -1, -1)

-        if self.refiner_vae:
-            xl = [x[:, :, :1, :, :]]
-            if x.shape[2] > self.ffactor_temporal:
-                xl += torch.split(x[:, :, 1: 1 + ((x.shape[2] - 1) // self.ffactor_temporal) * self.ffactor_temporal, :, :], self.ffactor_temporal * 2, dim=2)
-            x = xl
-        else:
-            x = [x]
-        out = []
+        x = self.conv_in(x)

-        conv_carry_in = None
+        for stage in self.down:
+            for blk in stage.block:
+                x = blk(x)
+            if hasattr(stage, 'downsample'):
+                x = stage.downsample(x)

-        for i, x1 in enumerate(x):
-            conv_carry_out = []
-            if i == len(x) - 1:
-                conv_carry_out = None
-            x1 = [ x1 ]
-            x1 = conv_carry_causal_3d(x1, self.conv_in, conv_carry_in, conv_carry_out)
-
-            for stage in self.down:
-                for blk in stage.block:
-                    x1 = blk(x1, conv_carry_in, conv_carry_out)
-                if hasattr(stage, 'downsample'):
-                    x1 = stage.downsample(x1, conv_carry_in, conv_carry_out)
-
-            out.append(x1)
-            conv_carry_in = conv_carry_out
-
-        if len(out) > 1:
-            out = torch.cat(out, dim=2)
-        else:
-            out = out[0]
-
-        x = self.mid.block_2(self.mid.attn_1(self.mid.block_1(out)))
-        del out
+        x = self.mid.block_2(self.mid.attn_1(self.mid.block_1(x)))

        b, c, t, h, w = x.shape
        grp = c // (self.z_channels << 1)
        skip = x.view(b, c // grp, grp, t, h, w).mean(2)

-        out = conv_carry_causal_3d([F.silu(self.norm_out(x))], self.conv_out) + skip
+        out = self.conv_out(F.silu(self.norm_out(x))) + skip

        if self.refiner_vae:
            out = self.regul(out)[0]

+            out = torch.cat((out[:, :, :1], out), dim=2)
+            out = out.permute(0, 2, 1, 3, 4)
+            b, f_times_2, c, h, w = out.shape
+            out = out.reshape(b, f_times_2 // 2, 2 * c, h, w)
+            out = out.permute(0, 2, 1, 3, 4).contiguous()
+
        return out

 class Decoder(nn.Module):
@@ -288,7 +239,7 @@ class Decoder(nn.Module):

        self.refiner_vae = refiner_vae
        if self.refiner_vae:
-            conv_op = NoPadConv3d
+            conv_op = VideoConv3d
            norm_op = RMS_norm
        else:
            conv_op = ops.Conv3d
@@ -298,9 +249,9 @@ class Decoder(nn.Module):
        self.conv_in = conv_op(z_channels, ch, kernel_size=3, stride=1, padding=1)

        self.mid = nn.Module()
-        self.mid.block_1 = HunyuanRefinerResnetBlock(in_channels=ch, out_channels=ch, conv_op=conv_op, norm_op=norm_op)
+        self.mid.block_1 = ResnetBlock(in_channels=ch, out_channels=ch, temb_channels=0, conv_op=conv_op, norm_op=norm_op)
        self.mid.attn_1 = AttnBlock(ch, conv_op=ops.Conv3d, norm_op=norm_op)
-        self.mid.block_2 = HunyuanRefinerResnetBlock(in_channels=ch, out_channels=ch,  conv_op=conv_op, norm_op=norm_op)
+        self.mid.block_2 = ResnetBlock(in_channels=ch, out_channels=ch, temb_channels=0, conv_op=conv_op, norm_op=norm_op)

        self.up = nn.ModuleList()
        depth = (ffactor_spatial >> 1).bit_length()
@@ -308,9 +259,10 @@ class Decoder(nn.Module):

        for i, tgt in enumerate(block_out_channels):
            stage = nn.Module()
-            stage.block = nn.ModuleList([HunyuanRefinerResnetBlock(in_channels=ch if j == 0 else tgt,
-                                                                   out_channels=tgt,
-                                                                   conv_op=conv_op, norm_op=norm_op)
+            stage.block = nn.ModuleList([ResnetBlock(in_channels=ch if j == 0 else tgt,
+                                                     out_channels=tgt,
+                                                     temb_channels=0,
+                                                     conv_op=conv_op, norm_op=norm_op)
                                        for j in range(num_res_blocks + 1)])
            ch = tgt
            if i < depth:
@@ -323,41 +275,27 @@ class Decoder(nn.Module):
        self.conv_out = conv_op(ch, out_channels, 3, stride=1, padding=1)

    def forward(self, z):
-        x = conv_carry_causal_3d([z], self.conv_in) + z.repeat_interleave(self.block_out_channels[0] // self.z_channels, 1)
+        if self.refiner_vae:
+            z = z.permute(0, 2, 1, 3, 4)
+            b, f, c, h, w = z.shape
+            z = z.reshape(b, f, 2, c // 2, h, w)
+            z = z.permute(0, 1, 2, 3, 4, 5).reshape(b, f * 2, c // 2, h, w)
+            z = z.permute(0, 2, 1, 3, 4)
+            z = z[:, :, 1:]
+
+        x = self.conv_in(z) + z.repeat_interleave(self.block_out_channels[0] // self.z_channels, 1)
        x = self.mid.block_2(self.mid.attn_1(self.mid.block_1(x)))

-        if self.refiner_vae:
-            x = torch.split(x, 2, dim=2)
-        else:
-            x = [ x ]
-        out = []
+        for stage in self.up:
+            for blk in stage.block:
+                x = blk(x)
+            if hasattr(stage, 'upsample'):
+                x = stage.upsample(x)

-        conv_carry_in = None
-
-        for i, x1 in enumerate(x):
-            conv_carry_out = []
-            if i == len(x) - 1:
-                conv_carry_out = None
-            for stage in self.up:
-                for blk in stage.block:
-                    x1 = blk(x1, conv_carry_in, conv_carry_out)
-                if hasattr(stage, 'upsample'):
-                    x1 = stage.upsample(x1, conv_carry_in, conv_carry_out)
-
-            x1 = [ F.silu(self.norm_out(x1)) ]
-            x1 = conv_carry_causal_3d(x1, self.conv_out, conv_carry_in, conv_carry_out)
-            out.append(x1)
-            conv_carry_in = conv_carry_out
-        del x
-
-        if len(out) > 1:
-            out = torch.cat(out, dim=2)
-        else:
-            out = out[0]
+        out = self.conv_out(F.silu(self.norm_out(x)))

        if not self.refiner_vae:
            if z.shape[-3] == 1:
                out = out[:, :, -1:]

        return out
-
--- a/comfy/ldm/lightricks/model.py
+++ b/comfy/ldm/lightricks/model.py
@@ -3,11 +3,12 @@ from torch import nn
 import comfy.patcher_extension
 import comfy.ldm.modules.attention
 import comfy.ldm.common_dit
+from einops import rearrange
 import math
 from typing import Dict, Optional, Tuple

 from .symmetric_patchifier import SymmetricPatchifier, latent_to_pixel_coords
-from comfy.ldm.flux.math import apply_rope1
+

 def get_timestep_embedding(
    timesteps: torch.Tensor,
@@ -237,6 +238,20 @@ class FeedForward(nn.Module):
        return self.net(x)


+def apply_rotary_emb(input_tensor, freqs_cis): #TODO: remove duplicate funcs and pick the best/fastest one
+    cos_freqs = freqs_cis[0]
+    sin_freqs = freqs_cis[1]
+
+    t_dup = rearrange(input_tensor, "... (d r) -> ... d r", r=2)
+    t1, t2 = t_dup.unbind(dim=-1)
+    t_dup = torch.stack((-t2, t1), dim=-1)
+    input_tensor_rot = rearrange(t_dup, "... d r -> ... (d r)")
+
+    out = input_tensor * cos_freqs + input_tensor_rot * sin_freqs
+
+    return out
+
+
 class CrossAttention(nn.Module):
    def __init__(self, query_dim, context_dim=None, heads=8, dim_head=64, dropout=0., attn_precision=None, dtype=None, device=None, operations=None):
        super().__init__()
@@ -266,8 +281,8 @@ class CrossAttention(nn.Module):
        k = self.k_norm(k)

        if pe is not None:
-            q = apply_rope1(q.unsqueeze(1), pe).squeeze(1)
-            k = apply_rope1(k.unsqueeze(1), pe).squeeze(1)
+            q = apply_rotary_emb(q, pe)
+            k = apply_rotary_emb(k, pe)

        if mask is None:
            out = comfy.ldm.modules.attention.optimized_attention(q, k, v, self.heads, attn_precision=self.attn_precision, transformer_options=transformer_options)
@@ -291,17 +306,12 @@ class BasicTransformerBlock(nn.Module):
    def forward(self, x, context=None, attention_mask=None, timestep=None, pe=None, transformer_options={}):
        shift_msa, scale_msa, gate_msa, shift_mlp, scale_mlp, gate_mlp = (self.scale_shift_table[None, None].to(device=x.device, dtype=x.dtype) + timestep.reshape(x.shape[0], timestep.shape[1], self.scale_shift_table.shape[0], -1)).unbind(dim=2)

-        attn1_input = comfy.ldm.common_dit.rms_norm(x)
-        attn1_input = torch.addcmul(attn1_input, attn1_input, scale_msa).add_(shift_msa)
-        attn1_input = self.attn1(attn1_input, pe=pe, transformer_options=transformer_options)
-        x.addcmul_(attn1_input, gate_msa)
-        del attn1_input
+        x += self.attn1(comfy.ldm.common_dit.rms_norm(x) * (1 + scale_msa) + shift_msa, pe=pe, transformer_options=transformer_options) * gate_msa

        x += self.attn2(x, context=context, mask=attention_mask, transformer_options=transformer_options)

-        y = comfy.ldm.common_dit.rms_norm(x)
-        y = torch.addcmul(y, y, scale_mlp).add_(shift_mlp)
-        x.addcmul_(self.ff(y), gate_mlp)
+        y = comfy.ldm.common_dit.rms_norm(x) * (1 + scale_mlp) + shift_mlp
+        x += self.ff(y) * gate_mlp

        return x

@@ -317,35 +327,41 @@ def get_fractional_positions(indices_grid, max_pos):


 def precompute_freqs_cis(indices_grid, dim, out_dtype, theta=10000.0, max_pos=[20, 2048, 2048]):
-    dtype = torch.float32
-    device = indices_grid.device
+    dtype = torch.float32 #self.dtype

-    # Get fractional positions and compute frequency indices
    fractional_positions = get_fractional_positions(indices_grid, max_pos)
-    indices = theta ** torch.linspace(0, 1, dim // 6, device=device, dtype=dtype) * math.pi / 2

-    # Compute frequencies and apply cos/sin
-    freqs = (indices * (fractional_positions.unsqueeze(-1) * 2 - 1)).transpose(-1, -2).flatten(2)
-    cos_vals = freqs.cos().repeat_interleave(2, dim=-1)
-    sin_vals = freqs.sin().repeat_interleave(2, dim=-1)
+    start = 1
+    end = theta
+    device = fractional_positions.device

-    # Pad if dim is not divisible by 6
+    indices = theta ** (
+        torch.linspace(
+            math.log(start, theta),
+            math.log(end, theta),
+            dim // 6,
+            device=device,
+            dtype=dtype,
+        )
+    )
+    indices = indices.to(dtype=dtype)
+
+    indices = indices * math.pi / 2
+
+    freqs = (
+        (indices * (fractional_positions.unsqueeze(-1) * 2 - 1))
+        .transpose(-1, -2)
+        .flatten(2)
+    )
+
+    cos_freq = freqs.cos().repeat_interleave(2, dim=-1)
+    sin_freq = freqs.sin().repeat_interleave(2, dim=-1)
    if dim % 6 != 0:
-        padding_size = dim % 6
-        cos_vals = torch.cat([torch.ones_like(cos_vals[:, :, :padding_size]), cos_vals], dim=-1)
-        sin_vals = torch.cat([torch.zeros_like(sin_vals[:, :, :padding_size]), sin_vals], dim=-1)
-
-    # Reshape and extract one value per pair (since repeat_interleave duplicates each value)
-    cos_vals = cos_vals.reshape(*cos_vals.shape[:2], -1, 2)[..., 0].to(out_dtype)  # [B, N, dim//2]
-    sin_vals = sin_vals.reshape(*sin_vals.shape[:2], -1, 2)[..., 0].to(out_dtype)  # [B, N, dim//2]
-
-    # Build rotation matrix [[cos, -sin], [sin, cos]] and add heads dimension
-    freqs_cis = torch.stack([
-        torch.stack([cos_vals, -sin_vals], dim=-1),
-        torch.stack([sin_vals, cos_vals], dim=-1)
-    ], dim=-2).unsqueeze(1)  # [B, 1, N, dim//2, 2, 2]
-
-    return freqs_cis
+        cos_padding = torch.ones_like(cos_freq[:, :, : dim % 6])
+        sin_padding = torch.zeros_like(cos_freq[:, :, : dim % 6])
+        cos_freq = torch.cat([cos_padding, cos_freq], dim=-1)
+        sin_freq = torch.cat([sin_padding, sin_freq], dim=-1)
+    return cos_freq.to(out_dtype), sin_freq.to(out_dtype)


 class LTXVModel(torch.nn.Module):
@@ -485,7 +501,7 @@ class LTXVModel(torch.nn.Module):
        shift, scale = scale_shift_values[:, :, 0], scale_shift_values[:, :, 1]
        x = self.norm_out(x)
        # Modulation
-        x = torch.addcmul(x, x, scale).add_(shift)
+        x = x * (1 + scale) + shift
        x = self.proj_out(x)

        x = self.patchifier.unpatchify(
--- a/comfy/ldm/lumina/model.py
+++ b/comfy/ldm/lumina/model.py
@@ -11,7 +11,6 @@ import comfy.ldm.common_dit
 from comfy.ldm.modules.diffusionmodules.mmdit import TimestepEmbedder
 from comfy.ldm.modules.attention import optimized_attention_masked
 from comfy.ldm.flux.layers import EmbedND
-from comfy.ldm.flux.math import apply_rope
 import comfy.patcher_extension


@@ -32,7 +31,6 @@ class JointAttention(nn.Module):
        n_heads: int,
        n_kv_heads: Optional[int],
        qk_norm: bool,
-        out_bias: bool = False,
        operation_settings={},
    ):
        """
@@ -61,7 +59,7 @@ class JointAttention(nn.Module):
        self.out = operation_settings.get("operations").Linear(
            n_heads * self.head_dim,
            dim,
-            bias=out_bias,
+            bias=False,
            device=operation_settings.get("device"),
            dtype=operation_settings.get("dtype"),
        )
@@ -72,6 +70,35 @@ class JointAttention(nn.Module):
        else:
            self.q_norm = self.k_norm = nn.Identity()

+    @staticmethod
+    def apply_rotary_emb(
+        x_in: torch.Tensor,
+        freqs_cis: torch.Tensor,
+    ) -> torch.Tensor:
+        """
+        Apply rotary embeddings to input tensors using the given frequency
+        tensor.
+
+        This function applies rotary embeddings to the given query 'xq' and
+        key 'xk' tensors using the provided frequency tensor 'freqs_cis'. The
+        input tensors are reshaped as complex numbers, and the frequency tensor
+        is reshaped for broadcasting compatibility. The resulting tensors
+        contain rotary embeddings and are returned as real tensors.
+
+        Args:
+            x_in (torch.Tensor): Query or Key tensor to apply rotary embeddings.
+            freqs_cis (torch.Tensor): Precomputed frequency tensor for complex
+                exponentials.
+
+        Returns:
+            Tuple[torch.Tensor, torch.Tensor]: Tuple of modified query tensor
+                and key tensor with rotary embeddings.
+        """
+
+        t_ = x_in.reshape(*x_in.shape[:-1], -1, 1, 2)
+        t_out = freqs_cis[..., 0] * t_[..., 0] + freqs_cis[..., 1] * t_[..., 1]
+        return t_out.reshape(*x_in.shape)
+
    def forward(
        self,
        x: torch.Tensor,
@@ -107,7 +134,8 @@ class JointAttention(nn.Module):
        xq = self.q_norm(xq)
        xk = self.k_norm(xk)

-        xq, xk = apply_rope(xq, xk, freqs_cis)
+        xq = JointAttention.apply_rotary_emb(xq, freqs_cis=freqs_cis)
+        xk = JointAttention.apply_rotary_emb(xk, freqs_cis=freqs_cis)

        n_rep = self.n_local_heads // self.n_local_kv_heads
        if n_rep >= 1:
@@ -187,8 +215,6 @@ class JointTransformerBlock(nn.Module):
        norm_eps: float,
        qk_norm: bool,
        modulation=True,
-        z_image_modulation=False,
-        attn_out_bias=False,
        operation_settings={},
    ) -> None:
        """
@@ -209,10 +235,10 @@ class JointTransformerBlock(nn.Module):
        super().__init__()
        self.dim = dim
        self.head_dim = dim // n_heads
-        self.attention = JointAttention(dim, n_heads, n_kv_heads, qk_norm, out_bias=attn_out_bias, operation_settings=operation_settings)
+        self.attention = JointAttention(dim, n_heads, n_kv_heads, qk_norm, operation_settings=operation_settings)
        self.feed_forward = FeedForward(
            dim=dim,
-            hidden_dim=dim,
+            hidden_dim=4 * dim,
            multiple_of=multiple_of,
            ffn_dim_multiplier=ffn_dim_multiplier,
            operation_settings=operation_settings,
@@ -226,27 +252,16 @@ class JointTransformerBlock(nn.Module):

        self.modulation = modulation
        if modulation:
-            if z_image_modulation:
-                self.adaLN_modulation = nn.Sequential(
-                    operation_settings.get("operations").Linear(
-                        min(dim, 256),
-                        4 * dim,
-                        bias=True,
-                        device=operation_settings.get("device"),
-                        dtype=operation_settings.get("dtype"),
-                    ),
-                )
-            else:
-                self.adaLN_modulation = nn.Sequential(
-                    nn.SiLU(),
-                    operation_settings.get("operations").Linear(
-                        min(dim, 1024),
-                        4 * dim,
-                        bias=True,
-                        device=operation_settings.get("device"),
-                        dtype=operation_settings.get("dtype"),
-                    ),
-                )
+            self.adaLN_modulation = nn.Sequential(
+                nn.SiLU(),
+                operation_settings.get("operations").Linear(
+                    min(dim, 1024),
+                    4 * dim,
+                    bias=True,
+                    device=operation_settings.get("device"),
+                    dtype=operation_settings.get("dtype"),
+                ),
+            )

    def forward(
        self,
@@ -308,7 +323,7 @@ class FinalLayer(nn.Module):
    The final layer of NextDiT.
    """

-    def __init__(self, hidden_size, patch_size, out_channels, z_image_modulation=False, operation_settings={}):
+    def __init__(self, hidden_size, patch_size, out_channels, operation_settings={}):
        super().__init__()
        self.norm_final = operation_settings.get("operations").LayerNorm(
            hidden_size,
@@ -325,15 +340,10 @@ class FinalLayer(nn.Module):
            dtype=operation_settings.get("dtype"),
        )

-        if z_image_modulation:
-            min_mod = 256
-        else:
-            min_mod = 1024
-
        self.adaLN_modulation = nn.Sequential(
            nn.SiLU(),
            operation_settings.get("operations").Linear(
-                min(hidden_size, min_mod),
+                min(hidden_size, 1024),
                hidden_size,
                bias=True,
                device=operation_settings.get("device"),
@@ -363,16 +373,12 @@ class NextDiT(nn.Module):
        n_heads: int = 32,
        n_kv_heads: Optional[int] = None,
        multiple_of: int = 256,
-        ffn_dim_multiplier: float = 4.0,
+        ffn_dim_multiplier: Optional[float] = None,
        norm_eps: float = 1e-5,
        qk_norm: bool = False,
        cap_feat_dim: int = 5120,
        axes_dims: List[int] = (16, 56, 56),
        axes_lens: List[int] = (1, 512, 512),
-        rope_theta=10000.0,
-        z_image_modulation=False,
-        time_scale=1.0,
-        pad_tokens_multiple=None,
        image_model=None,
        device=None,
        dtype=None,
@@ -384,8 +390,6 @@ class NextDiT(nn.Module):
        self.in_channels = in_channels
        self.out_channels = in_channels
        self.patch_size = patch_size
-        self.time_scale = time_scale
-        self.pad_tokens_multiple = pad_tokens_multiple

        self.x_embedder = operation_settings.get("operations").Linear(
            in_features=patch_size * patch_size * in_channels,
@@ -407,7 +411,6 @@ class NextDiT(nn.Module):
                    norm_eps,
                    qk_norm,
                    modulation=True,
-                    z_image_modulation=z_image_modulation,
                    operation_settings=operation_settings,
                )
                for layer_id in range(n_refiner_layers)
@@ -431,7 +434,7 @@ class NextDiT(nn.Module):
            ]
        )

-        self.t_embedder = TimestepEmbedder(min(dim, 1024), output_size=256 if z_image_modulation else None, **operation_settings)
+        self.t_embedder = TimestepEmbedder(min(dim, 1024), **operation_settings)
        self.cap_embedder = nn.Sequential(
            operation_settings.get("operations").RMSNorm(cap_feat_dim, eps=norm_eps, elementwise_affine=True, device=operation_settings.get("device"), dtype=operation_settings.get("dtype")),
            operation_settings.get("operations").Linear(
@@ -454,24 +457,18 @@ class NextDiT(nn.Module):
                    ffn_dim_multiplier,
                    norm_eps,
                    qk_norm,
-                    z_image_modulation=z_image_modulation,
-                    attn_out_bias=False,
                    operation_settings=operation_settings,
                )
                for layer_id in range(n_layers)
            ]
        )
        self.norm_final = operation_settings.get("operations").RMSNorm(dim, eps=norm_eps, elementwise_affine=True, device=operation_settings.get("device"), dtype=operation_settings.get("dtype"))
-        self.final_layer = FinalLayer(dim, patch_size, self.out_channels, z_image_modulation=z_image_modulation, operation_settings=operation_settings)
-
-        if self.pad_tokens_multiple is not None:
-            self.x_pad_token = nn.Parameter(torch.empty((1, dim), device=device, dtype=dtype))
-            self.cap_pad_token = nn.Parameter(torch.empty((1, dim), device=device, dtype=dtype))
+        self.final_layer = FinalLayer(dim, patch_size, self.out_channels, operation_settings=operation_settings)

        assert (dim // n_heads) == sum(axes_dims)
        self.axes_dims = axes_dims
        self.axes_lens = axes_lens
-        self.rope_embedder = EmbedND(dim=dim // n_heads, theta=rope_theta, axes_dim=axes_dims)
+        self.rope_embedder = EmbedND(dim=dim // n_heads, theta=10000.0, axes_dim=axes_dims)
        self.dim = dim
        self.n_heads = n_heads

@@ -506,54 +503,96 @@ class NextDiT(nn.Module):
        bsz = len(x)
        pH = pW = self.patch_size
        device = x[0].device
+        dtype = x[0].dtype

-        if self.pad_tokens_multiple is not None:
-            pad_extra = (-cap_feats.shape[1]) % self.pad_tokens_multiple
-            cap_feats = torch.cat((cap_feats, self.cap_pad_token.to(device=cap_feats.device, dtype=cap_feats.dtype, copy=True).unsqueeze(0).repeat(cap_feats.shape[0], pad_extra, 1)), dim=1)
+        if cap_mask is not None:
+            l_effective_cap_len = cap_mask.sum(dim=1).tolist()
+        else:
+            l_effective_cap_len = [num_tokens] * bsz

-        cap_pos_ids = torch.zeros(bsz, cap_feats.shape[1], 3, dtype=torch.float32, device=device)
-        cap_pos_ids[:, :, 0] = torch.arange(cap_feats.shape[1], dtype=torch.float32, device=device) + 1.0
+        if cap_mask is not None and not torch.is_floating_point(cap_mask):
+            cap_mask = (cap_mask - 1).to(dtype) * torch.finfo(dtype).max

-        B, C, H, W = x.shape
-        x = self.x_embedder(x.view(B, C, H // pH, pH, W // pW, pW).permute(0, 2, 4, 3, 5, 1).flatten(3).flatten(1, 2))
+        img_sizes = [(img.size(1), img.size(2)) for img in x]
+        l_effective_img_len = [(H // pH) * (W // pW) for (H, W) in img_sizes]

-        rope_options = transformer_options.get("rope_options", None)
-        h_scale = 1.0
-        w_scale = 1.0
-        h_start = 0
-        w_start = 0
-        if rope_options is not None:
-            h_scale = rope_options.get("scale_y", 1.0)
-            w_scale = rope_options.get("scale_x", 1.0)
+        max_seq_len = max(
+            (cap_len+img_len for cap_len, img_len in zip(l_effective_cap_len, l_effective_img_len))
+        )
+        max_cap_len = max(l_effective_cap_len)
+        max_img_len = max(l_effective_img_len)

-            h_start = rope_options.get("shift_y", 0.0)
-            w_start = rope_options.get("shift_x", 0.0)
+        position_ids = torch.zeros(bsz, max_seq_len, 3, dtype=torch.int32, device=device)

-        H_tokens, W_tokens = H // pH, W // pW
-        x_pos_ids = torch.zeros((bsz, x.shape[1], 3), dtype=torch.float32, device=device)
-        x_pos_ids[:, :, 0] = cap_feats.shape[1] + 1
-        x_pos_ids[:, :, 1] = (torch.arange(H_tokens, dtype=torch.float32, device=device) * h_scale + h_start).view(-1, 1).repeat(1, W_tokens).flatten()
-        x_pos_ids[:, :, 2] = (torch.arange(W_tokens, dtype=torch.float32, device=device) * w_scale + w_start).view(1, -1).repeat(H_tokens, 1).flatten()
+        for i in range(bsz):
+            cap_len = l_effective_cap_len[i]
+            img_len = l_effective_img_len[i]
+            H, W = img_sizes[i]
+            H_tokens, W_tokens = H // pH, W // pW
+            assert H_tokens * W_tokens == img_len

-        if self.pad_tokens_multiple is not None:
-            pad_extra = (-x.shape[1]) % self.pad_tokens_multiple
-            x = torch.cat((x, self.x_pad_token.to(device=x.device, dtype=x.dtype, copy=True).unsqueeze(0).repeat(x.shape[0], pad_extra, 1)), dim=1)
-            x_pos_ids = torch.nn.functional.pad(x_pos_ids, (0, 0, 0, pad_extra))
+            position_ids[i, :cap_len, 0] = torch.arange(cap_len, dtype=torch.int32, device=device)
+            position_ids[i, cap_len:cap_len+img_len, 0] = cap_len
+            row_ids = torch.arange(H_tokens, dtype=torch.int32, device=device).view(-1, 1).repeat(1, W_tokens).flatten()
+            col_ids = torch.arange(W_tokens, dtype=torch.int32, device=device).view(1, -1).repeat(H_tokens, 1).flatten()
+            position_ids[i, cap_len:cap_len+img_len, 1] = row_ids
+            position_ids[i, cap_len:cap_len+img_len, 2] = col_ids

-        freqs_cis = self.rope_embedder(torch.cat((cap_pos_ids, x_pos_ids), dim=1)).movedim(1, 2)
+        freqs_cis = self.rope_embedder(position_ids).movedim(1, 2).to(dtype)
+
+        # build freqs_cis for cap and image individually
+        cap_freqs_cis_shape = list(freqs_cis.shape)
+        # cap_freqs_cis_shape[1] = max_cap_len
+        cap_freqs_cis_shape[1] = cap_feats.shape[1]
+        cap_freqs_cis = torch.zeros(*cap_freqs_cis_shape, device=device, dtype=freqs_cis.dtype)
+
+        img_freqs_cis_shape = list(freqs_cis.shape)
+        img_freqs_cis_shape[1] = max_img_len
+        img_freqs_cis = torch.zeros(*img_freqs_cis_shape, device=device, dtype=freqs_cis.dtype)
+
+        for i in range(bsz):
+            cap_len = l_effective_cap_len[i]
+            img_len = l_effective_img_len[i]
+            cap_freqs_cis[i, :cap_len] = freqs_cis[i, :cap_len]
+            img_freqs_cis[i, :img_len] = freqs_cis[i, cap_len:cap_len+img_len]

        # refine context
        for layer in self.context_refiner:
-            cap_feats = layer(cap_feats, cap_mask, freqs_cis[:, :cap_pos_ids.shape[1]], transformer_options=transformer_options)
+            cap_feats = layer(cap_feats, cap_mask, cap_freqs_cis, transformer_options=transformer_options)

-        padded_img_mask = None
+        # refine image
+        flat_x = []
+        for i in range(bsz):
+            img = x[i]
+            C, H, W = img.size()
+            img = img.view(C, H // pH, pH, W // pW, pW).permute(1, 3, 2, 4, 0).flatten(2).flatten(0, 1)
+            flat_x.append(img)
+        x = flat_x
+        padded_img_embed = torch.zeros(bsz, max_img_len, x[0].shape[-1], device=device, dtype=x[0].dtype)
+        padded_img_mask = torch.zeros(bsz, max_img_len, dtype=dtype, device=device)
+        for i in range(bsz):
+            padded_img_embed[i, :l_effective_img_len[i]] = x[i]
+            padded_img_mask[i, l_effective_img_len[i]:] = -torch.finfo(dtype).max
+
+        padded_img_embed = self.x_embedder(padded_img_embed)
+        padded_img_mask = padded_img_mask.unsqueeze(1)
        for layer in self.noise_refiner:
-            x = layer(x, padded_img_mask, freqs_cis[:, cap_pos_ids.shape[1]:], t, transformer_options=transformer_options)
+            padded_img_embed = layer(padded_img_embed, padded_img_mask, img_freqs_cis, t, transformer_options=transformer_options)
+
+        if cap_mask is not None:
+            mask = torch.zeros(bsz, max_seq_len, dtype=dtype, device=device)
+            mask[:, :max_cap_len] = cap_mask[:, :max_cap_len]
+        else:
+            mask = None
+
+        padded_full_embed = torch.zeros(bsz, max_seq_len, self.dim, device=device, dtype=x[0].dtype)
+        for i in range(bsz):
+            cap_len = l_effective_cap_len[i]
+            img_len = l_effective_img_len[i]
+
+            padded_full_embed[i, :cap_len] = cap_feats[i, :cap_len]
+            padded_full_embed[i, cap_len:cap_len+img_len] = padded_img_embed[i, :img_len]

-        padded_full_embed = torch.cat((cap_feats, x), dim=1)
-        mask = None
-        img_sizes = [(H, W)] * bsz
-        l_effective_cap_len = [cap_feats.shape[1]] * bsz
        return padded_full_embed, mask, img_sizes, l_effective_cap_len, freqs_cis

    def forward(self, x, timesteps, context, num_tokens, attention_mask=None, **kwargs):
@@ -576,7 +615,7 @@ class NextDiT(nn.Module):
        y: (N,) tensor of text tokens/features
        """

-        t = self.t_embedder(t * self.time_scale, dtype=x.dtype)  # (N, D)
+        t = self.t_embedder(t, dtype=x.dtype)  # (N, D)
        adaln_input = t

        cap_feats = self.cap_embedder(cap_feats)  # (N, L, D)  # todo check if able to batchify w.o. redundant compute
--- a/comfy/ldm/models/autoencoder.py
+++ b/comfy/ldm/models/autoencoder.py
@@ -9,8 +9,6 @@ from comfy.ldm.modules.distributions.distributions import DiagonalGaussianDistri
 from comfy.ldm.util import get_obj_from_str, instantiate_from_config
 from comfy.ldm.modules.ema import LitEma
 import comfy.ops
-from einops import rearrange
-import comfy.model_management

 class DiagonalGaussianRegularizer(torch.nn.Module):
    def __init__(self, sample: bool = False):
@@ -181,21 +179,6 @@ class AutoencodingEngineLegacy(AutoencodingEngine):
        self.post_quant_conv = conv_op(embed_dim, ddconfig["z_channels"], 1)
        self.embed_dim = embed_dim

-        if ddconfig.get("batch_norm_latent", False):
-            self.bn_eps = 1e-4
-            self.bn_momentum = 0.1
-            self.ps = [2, 2]
-            self.bn = torch.nn.BatchNorm2d(math.prod(self.ps) * ddconfig["z_channels"],
-                                           eps=self.bn_eps,
-                                           momentum=self.bn_momentum,
-                                           affine=False,
-                                           track_running_stats=True,
-                                           )
-            self.bn.eval()
-        else:
-            self.bn = None
-
-
    def get_autoencoder_params(self) -> list:
        params = super().get_autoencoder_params()
        return params
@@ -218,36 +201,11 @@ class AutoencodingEngineLegacy(AutoencodingEngine):
            z = torch.cat(z, 0)

        z, reg_log = self.regularization(z)
-
-        if self.bn is not None:
-            z = rearrange(z,
-                          "... c (i pi) (j pj)  -> ... (c pi pj) i j",
-                          pi=self.ps[0],
-                          pj=self.ps[1],
-                          )
-
-            z = torch.nn.functional.batch_norm(z,
-                                               comfy.model_management.cast_to(self.bn.running_mean, dtype=z.dtype, device=z.device),
-                                               comfy.model_management.cast_to(self.bn.running_var, dtype=z.dtype, device=z.device),
-                                               momentum=self.bn_momentum,
-                                               eps=self.bn_eps)
-
        if return_reg_log:
            return z, reg_log
        return z

    def decode(self, z: torch.Tensor, **decoder_kwargs) -> torch.Tensor:
-        if self.bn is not None:
-            s = torch.sqrt(comfy.model_management.cast_to(self.bn.running_var.view(1, -1, 1, 1), dtype=z.dtype, device=z.device) + self.bn_eps)
-            m = comfy.model_management.cast_to(self.bn.running_mean.view(1, -1, 1, 1), dtype=z.dtype, device=z.device)
-            z = z * s + m
-            z = rearrange(
-                z,
-                "... (c pi pj) i j -> ... c (i pi) (j pj)",
-                pi=self.ps[0],
-                pj=self.ps[1],
-            )
-
        if self.max_batch_size is None:
            dec = self.post_quant_conv(z)
            dec = self.decoder(dec, **decoder_kwargs)
--- a/comfy/ldm/modules/diffusionmodules/mmdit.py
+++ b/comfy/ldm/modules/diffusionmodules/mmdit.py
@@ -211,14 +211,12 @@ class TimestepEmbedder(nn.Module):
    Embeds scalar timesteps into vector representations.
    """

-    def __init__(self, hidden_size, frequency_embedding_size=256, output_size=None, dtype=None, device=None, operations=None):
+    def __init__(self, hidden_size, frequency_embedding_size=256, dtype=None, device=None, operations=None):
        super().__init__()
-        if output_size is None:
-            output_size = hidden_size
        self.mlp = nn.Sequential(
            operations.Linear(frequency_embedding_size, hidden_size, bias=True, dtype=dtype, device=device),
            nn.SiLU(),
-            operations.Linear(hidden_size, output_size, bias=True, dtype=dtype, device=device),
+            operations.Linear(hidden_size, hidden_size, bias=True, dtype=dtype, device=device),
        )
        self.frequency_embedding_size = frequency_embedding_size

--- a/comfy/ldm/qwen_image/controlnet.py
+++ b/comfy/ldm/qwen_image/controlnet.py
@@ -44,7 +44,7 @@ class QwenImageControlNetModel(QwenImageTransformer2DModel):
        txt_start = round(max(((x.shape[-1] + (self.patch_size // 2)) // self.patch_size) // 2, ((x.shape[-2] + (self.patch_size // 2)) // self.patch_size) // 2))
        txt_ids = torch.arange(txt_start, txt_start + context.shape[1], device=x.device).reshape(1, -1, 1).repeat(x.shape[0], 1, 3)
        ids = torch.cat((txt_ids, img_ids), dim=1)
-        image_rotary_emb = self.pe_embedder(ids).to(x.dtype).contiguous()
+        image_rotary_emb = self.pe_embedder(ids).squeeze(1).unsqueeze(2).to(x.dtype)
        del ids, txt_ids, img_ids

        hidden_states = self.img_in(hidden_states) + self.controlnet_x_embedder(hint)
--- a/comfy/ldm/qwen_image/model.py
+++ b/comfy/ldm/qwen_image/model.py
@@ -5,12 +5,13 @@ import torch.nn.functional as F
 from typing import Optional, Tuple
 from einops import repeat

+from comfy.ldm.flipflop_transformer import FlipFlopModule
 from comfy.ldm.lightricks.model import TimestepEmbedding, Timesteps
 from comfy.ldm.modules.attention import optimized_attention_masked
 from comfy.ldm.flux.layers import EmbedND
 import comfy.ldm.common_dit
 import comfy.patcher_extension
-from comfy.ldm.flux.math import apply_rope1
+import comfy.ops

 class GELU(nn.Module):
    def __init__(self, dim_in: int, dim_out: int, approximate: str = "none", bias: bool = True, dtype=None, device=None, operations=None):
@@ -135,34 +136,33 @@ class Attention(nn.Module):
        image_rotary_emb: Optional[torch.Tensor] = None,
        transformer_options={},
    ) -> Tuple[torch.Tensor, torch.Tensor]:
-        batch_size = hidden_states.shape[0]
-        seq_img = hidden_states.shape[1]
        seq_txt = encoder_hidden_states.shape[1]

-        # Project and reshape to BHND format (batch, heads, seq, dim)
-        img_query = self.to_q(hidden_states).view(batch_size, seq_img, self.heads, -1).transpose(1, 2).contiguous()
-        img_key = self.to_k(hidden_states).view(batch_size, seq_img, self.heads, -1).transpose(1, 2).contiguous()
-        img_value = self.to_v(hidden_states).view(batch_size, seq_img, self.heads, -1).transpose(1, 2)
+        img_query = self.to_q(hidden_states).unflatten(-1, (self.heads, -1))
+        img_key = self.to_k(hidden_states).unflatten(-1, (self.heads, -1))
+        img_value = self.to_v(hidden_states).unflatten(-1, (self.heads, -1))

-        txt_query = self.add_q_proj(encoder_hidden_states).view(batch_size, seq_txt, self.heads, -1).transpose(1, 2).contiguous()
-        txt_key = self.add_k_proj(encoder_hidden_states).view(batch_size, seq_txt, self.heads, -1).transpose(1, 2).contiguous()
-        txt_value = self.add_v_proj(encoder_hidden_states).view(batch_size, seq_txt, self.heads, -1).transpose(1, 2)
+        txt_query = self.add_q_proj(encoder_hidden_states).unflatten(-1, (self.heads, -1))
+        txt_key = self.add_k_proj(encoder_hidden_states).unflatten(-1, (self.heads, -1))
+        txt_value = self.add_v_proj(encoder_hidden_states).unflatten(-1, (self.heads, -1))

        img_query = self.norm_q(img_query)
        img_key = self.norm_k(img_key)
        txt_query = self.norm_added_q(txt_query)
        txt_key = self.norm_added_k(txt_key)

-        joint_query = torch.cat([txt_query, img_query], dim=2)
-        joint_key = torch.cat([txt_key, img_key], dim=2)
-        joint_value = torch.cat([txt_value, img_value], dim=2)
+        joint_query = torch.cat([txt_query, img_query], dim=1)
+        joint_key = torch.cat([txt_key, img_key], dim=1)
+        joint_value = torch.cat([txt_value, img_value], dim=1)

-        joint_query = apply_rope1(joint_query, image_rotary_emb)
-        joint_key = apply_rope1(joint_key, image_rotary_emb)
+        joint_query = apply_rotary_emb(joint_query, image_rotary_emb)
+        joint_key = apply_rotary_emb(joint_key, image_rotary_emb)

-        joint_hidden_states = optimized_attention_masked(joint_query, joint_key, joint_value, self.heads,
-                                                         attention_mask, transformer_options=transformer_options,
-                                                         skip_reshape=True)
+        joint_query = joint_query.flatten(start_dim=2)
+        joint_key = joint_key.flatten(start_dim=2)
+        joint_value = joint_value.flatten(start_dim=2)
+
+        joint_hidden_states = optimized_attention_masked(joint_query, joint_key, joint_value, self.heads, attention_mask, transformer_options=transformer_options)

        txt_attn_output = joint_hidden_states[:, :seq_txt, :]
        img_attn_output = joint_hidden_states[:, seq_txt:, :]
@@ -236,10 +236,10 @@ class QwenImageTransformerBlock(nn.Module):
        img_mod1, img_mod2 = img_mod_params.chunk(2, dim=-1)
        txt_mod1, txt_mod2 = txt_mod_params.chunk(2, dim=-1)

-        img_modulated, img_gate1 = self._modulate(self.img_norm1(hidden_states), img_mod1)
-        del img_mod1
-        txt_modulated, txt_gate1 = self._modulate(self.txt_norm1(encoder_hidden_states), txt_mod1)
-        del txt_mod1
+        img_normed = self.img_norm1(hidden_states)
+        img_modulated, img_gate1 = self._modulate(img_normed, img_mod1)
+        txt_normed = self.txt_norm1(encoder_hidden_states)
+        txt_modulated, txt_gate1 = self._modulate(txt_normed, txt_mod1)

        img_attn_output, txt_attn_output = self.attn(
            hidden_states=img_modulated,
@@ -248,20 +248,16 @@ class QwenImageTransformerBlock(nn.Module):
            image_rotary_emb=image_rotary_emb,
            transformer_options=transformer_options,
        )
-        del img_modulated
-        del txt_modulated

        hidden_states = hidden_states + img_gate1 * img_attn_output
        encoder_hidden_states = encoder_hidden_states + txt_gate1 * txt_attn_output
-        del img_attn_output
-        del txt_attn_output
-        del img_gate1
-        del txt_gate1

-        img_modulated2, img_gate2 = self._modulate(self.img_norm2(hidden_states), img_mod2)
+        img_normed2 = self.img_norm2(hidden_states)
+        img_modulated2, img_gate2 = self._modulate(img_normed2, img_mod2)
        hidden_states = torch.addcmul(hidden_states, img_gate2, self.img_mlp(img_modulated2))

-        txt_modulated2, txt_gate2 = self._modulate(self.txt_norm2(encoder_hidden_states), txt_mod2)
+        txt_normed2 = self.txt_norm2(encoder_hidden_states)
+        txt_modulated2, txt_gate2 = self._modulate(txt_normed2, txt_mod2)
        encoder_hidden_states = torch.addcmul(encoder_hidden_states, txt_gate2, self.txt_mlp(txt_modulated2))

        return encoder_hidden_states, hidden_states
@@ -289,7 +285,7 @@ class LastLayer(nn.Module):
        return x


-class QwenImageTransformer2DModel(nn.Module):
+class QwenImageTransformer2DModel(FlipFlopModule):
    def __init__(
        self,
        patch_size: int = 2,
@@ -306,9 +302,9 @@ class QwenImageTransformer2DModel(nn.Module):
        final_layer=True,
        dtype=None,
        device=None,
-        operations=None,
+        operations: comfy.ops.disable_weight_init=None,
    ):
-        super().__init__()
+        super().__init__(block_types=("transformer_blocks",))
        self.dtype = dtype
        self.patch_size = patch_size
        self.in_channels = in_channels
@@ -372,6 +368,40 @@ class QwenImageTransformer2DModel(nn.Module):
            comfy.patcher_extension.get_all_wrappers(comfy.patcher_extension.WrappersMP.DIFFUSION_MODEL, transformer_options)
        ).execute(x, timestep, context, attention_mask, guidance, ref_latents, transformer_options, **kwargs)

+    def indiv_block_fwd(self, i, block, hidden_states, encoder_hidden_states, encoder_hidden_states_mask, temb, image_rotary_emb, patches, control, blocks_replace, x, transformer_options):
+        if ("double_block", i) in blocks_replace:
+            def block_wrap(args):
+                out = {}
+                out["txt"], out["img"] = block(hidden_states=args["img"], encoder_hidden_states=args["txt"], encoder_hidden_states_mask=encoder_hidden_states_mask, temb=args["vec"], image_rotary_emb=args["pe"], transformer_options=args["transformer_options"])
+                return out
+            out = blocks_replace[("double_block", i)]({"img": hidden_states, "txt": encoder_hidden_states, "vec": temb, "pe": image_rotary_emb, "transformer_options": transformer_options}, {"original_block": block_wrap})
+            hidden_states = out["img"]
+            encoder_hidden_states = out["txt"]
+        else:
+            encoder_hidden_states, hidden_states = block(
+                hidden_states=hidden_states,
+                encoder_hidden_states=encoder_hidden_states,
+                encoder_hidden_states_mask=encoder_hidden_states_mask,
+                temb=temb,
+                image_rotary_emb=image_rotary_emb,
+                transformer_options=transformer_options,
+            )
+
+        if "double_block" in patches:
+            for p in patches["double_block"]:
+                out = p({"img": hidden_states, "txt": encoder_hidden_states, "x": x, "block_index": i, "transformer_options": transformer_options})
+                hidden_states = out["img"]
+                encoder_hidden_states = out["txt"]
+
+        if control is not None: # Controlnet
+            control_i = control.get("input")
+            if i < len(control_i):
+                add = control_i[i]
+                if add is not None:
+                    hidden_states[:, :add.shape[1]] += add
+
+        return hidden_states, encoder_hidden_states
+
    def _forward(
        self,
        x,
@@ -419,7 +449,7 @@ class QwenImageTransformer2DModel(nn.Module):
        txt_start = round(max(((x.shape[-1] + (self.patch_size // 2)) // self.patch_size) // 2, ((x.shape[-2] + (self.patch_size // 2)) // self.patch_size) // 2))
        txt_ids = torch.arange(txt_start, txt_start + context.shape[1], device=x.device).reshape(1, -1, 1).repeat(x.shape[0], 1, 3)
        ids = torch.cat((txt_ids, img_ids), dim=1)
-        image_rotary_emb = self.pe_embedder(ids).to(x.dtype).contiguous()
+        image_rotary_emb = self.pe_embedder(ids).squeeze(1).unsqueeze(2).to(x.dtype)
        del ids, txt_ids, img_ids

        hidden_states = self.img_in(hidden_states)
@@ -439,40 +469,8 @@ class QwenImageTransformer2DModel(nn.Module):
        patches = transformer_options.get("patches", {})
        blocks_replace = patches_replace.get("dit", {})

-        transformer_options["total_blocks"] = len(self.transformer_blocks)
-        transformer_options["block_type"] = "double"
-        for i, block in enumerate(self.transformer_blocks):
-            transformer_options["block_index"] = i
-            if ("double_block", i) in blocks_replace:
-                def block_wrap(args):
-                    out = {}
-                    out["txt"], out["img"] = block(hidden_states=args["img"], encoder_hidden_states=args["txt"], encoder_hidden_states_mask=encoder_hidden_states_mask, temb=args["vec"], image_rotary_emb=args["pe"], transformer_options=args["transformer_options"])
-                    return out
-                out = blocks_replace[("double_block", i)]({"img": hidden_states, "txt": encoder_hidden_states, "vec": temb, "pe": image_rotary_emb, "transformer_options": transformer_options}, {"original_block": block_wrap})
-                hidden_states = out["img"]
-                encoder_hidden_states = out["txt"]
-            else:
-                encoder_hidden_states, hidden_states = block(
-                    hidden_states=hidden_states,
-                    encoder_hidden_states=encoder_hidden_states,
-                    encoder_hidden_states_mask=encoder_hidden_states_mask,
-                    temb=temb,
-                    image_rotary_emb=image_rotary_emb,
-                    transformer_options=transformer_options,
-                )
-
-            if "double_block" in patches:
-                for p in patches["double_block"]:
-                    out = p({"img": hidden_states, "txt": encoder_hidden_states, "x": x, "block_index": i, "transformer_options": transformer_options})
-                    hidden_states = out["img"]
-                    encoder_hidden_states = out["txt"]
-
-            if control is not None: # Controlnet
-                control_i = control.get("input")
-                if i < len(control_i):
-                    add = control_i[i]
-                    if add is not None:
-                        hidden_states[:, :add.shape[1]] += add
+        out = (hidden_states, encoder_hidden_states)
+        hidden_states, encoder_hidden_states = self.execute_blocks("transformer_blocks", self.indiv_block_fwd, out, encoder_hidden_states_mask, temb, image_rotary_emb, patches, control, blocks_replace, x, transformer_options)

        hidden_states = self.norm_out(hidden_states, temb)
        hidden_states = self.proj_out(hidden_states)
--- a/comfy/ldm/wan/model.py
+++ b/comfy/ldm/wan/model.py
@@ -7,6 +7,7 @@ import torch.nn as nn
 from einops import rearrange

 from comfy.ldm.modules.attention import optimized_attention
+from comfy.ldm.flipflop_transformer import FlipFlopModule
 from comfy.ldm.flux.layers import EmbedND
 from comfy.ldm.flux.math import apply_rope1
 import comfy.ldm.common_dit
@@ -232,7 +233,6 @@ class WanAttentionBlock(nn.Module):
        # assert e[0].dtype == torch.float32

        # self-attention
-        x = x.contiguous() # otherwise implicit in LayerNorm
        y = self.self_attn(
            torch.addcmul(repeat_e(e[0], x), self.norm1(x), 1 + repeat_e(e[1], x)),
            freqs, transformer_options=transformer_options)
@@ -385,7 +385,7 @@ class MLPProj(torch.nn.Module):
        return clip_extra_context_tokens


-class WanModel(torch.nn.Module):
+class WanModel(FlipFlopModule):
    r"""
    Wan diffusion backbone supporting both text-to-video and image-to-video.
    """
@@ -413,6 +413,7 @@ class WanModel(torch.nn.Module):
                 device=None,
                 dtype=None,
                 operations=None,
+                 enable_flipflop=True,
                 ):
        r"""
        Initialize the diffusion model backbone.
@@ -450,7 +451,7 @@ class WanModel(torch.nn.Module):
                Epsilon value for normalization layers
        """

-        super().__init__()
+        super().__init__(block_types=("blocks",), enable_flipflop=enable_flipflop)
        self.dtype = dtype
        operation_settings = {"operations": operations, "device": device, "dtype": dtype}

@@ -507,6 +508,18 @@ class WanModel(torch.nn.Module):
        else:
            self.ref_conv = None

+    def indiv_block_fwd(self, i, block, x, e0, freqs, context, context_img_len, blocks_replace, transformer_options):
+        if ("double_block", i) in blocks_replace:
+            def block_wrap(args):
+                out = {}
+                out["img"] = block(args["img"], context=args["txt"], e=args["vec"], freqs=args["pe"], context_img_len=context_img_len, transformer_options=args["transformer_options"])
+                return out
+            out = blocks_replace[("double_block", i)]({"img": x, "txt": context, "vec": e0, "pe": freqs, "transformer_options": transformer_options}, {"original_block": block_wrap})
+            x = out["img"]
+        else:
+            x = block(x, e=e0, freqs=freqs, context=context, context_img_len=context_img_len, transformer_options=transformer_options)
+        return x
+
    def forward_orig(
        self,
        x,
@@ -568,16 +581,8 @@ class WanModel(torch.nn.Module):

        patches_replace = transformer_options.get("patches_replace", {})
        blocks_replace = patches_replace.get("dit", {})
-        for i, block in enumerate(self.blocks):
-            if ("double_block", i) in blocks_replace:
-                def block_wrap(args):
-                    out = {}
-                    out["img"] = block(args["img"], context=args["txt"], e=args["vec"], freqs=args["pe"], context_img_len=context_img_len, transformer_options=args["transformer_options"])
-                    return out
-                out = blocks_replace[("double_block", i)]({"img": x, "txt": context, "vec": e0, "pe": freqs, "transformer_options": transformer_options}, {"original_block": block_wrap})
-                x = out["img"]
-            else:
-                x = block(x, e=e0, freqs=freqs, context=context, context_img_len=context_img_len, transformer_options=transformer_options)
+        # execute blocks
+        x = self.execute_blocks("blocks", self.indiv_block_fwd, x, e0, freqs, context, context_img_len, blocks_replace, transformer_options)

        # head
        x = self.head(x, e)
@@ -589,7 +594,7 @@ class WanModel(torch.nn.Module):
        x = self.unpatchify(x, grid_sizes)
        return x

-    def rope_encode(self, t, h, w, t_start=0, steps_t=None, steps_h=None, steps_w=None, device=None, dtype=None, transformer_options={}):
+    def rope_encode(self, t, h, w, t_start=0, steps_t=None, steps_h=None, steps_w=None, device=None, dtype=None):
        patch_size = self.patch_size
        t_len = ((t + (patch_size[0] // 2)) // patch_size[0])
        h_len = ((h + (patch_size[1] // 2)) // patch_size[1])
@@ -602,22 +607,10 @@ class WanModel(torch.nn.Module):
        if steps_w is None:
            steps_w = w_len

-        h_start = 0
-        w_start = 0
-        rope_options = transformer_options.get("rope_options", None)
-        if rope_options is not None:
-            t_len = (t_len - 1.0) * rope_options.get("scale_t", 1.0) + 1.0
-            h_len = (h_len - 1.0) * rope_options.get("scale_y", 1.0) + 1.0
-            w_len = (w_len - 1.0) * rope_options.get("scale_x", 1.0) + 1.0
-
-            t_start += rope_options.get("shift_t", 0.0)
-            h_start += rope_options.get("shift_y", 0.0)
-            w_start += rope_options.get("shift_x", 0.0)
-
        img_ids = torch.zeros((steps_t, steps_h, steps_w, 3), device=device, dtype=dtype)
        img_ids[:, :, :, 0] = img_ids[:, :, :, 0] + torch.linspace(t_start, t_start + (t_len - 1), steps=steps_t, device=device, dtype=dtype).reshape(-1, 1, 1)
-        img_ids[:, :, :, 1] = img_ids[:, :, :, 1] + torch.linspace(h_start, h_start + (h_len - 1), steps=steps_h, device=device, dtype=dtype).reshape(1, -1, 1)
-        img_ids[:, :, :, 2] = img_ids[:, :, :, 2] + torch.linspace(w_start, w_start + (w_len - 1), steps=steps_w, device=device, dtype=dtype).reshape(1, 1, -1)
+        img_ids[:, :, :, 1] = img_ids[:, :, :, 1] + torch.linspace(0, h_len - 1, steps=steps_h, device=device, dtype=dtype).reshape(1, -1, 1)
+        img_ids[:, :, :, 2] = img_ids[:, :, :, 2] + torch.linspace(0, w_len - 1, steps=steps_w, device=device, dtype=dtype).reshape(1, 1, -1)
        img_ids = img_ids.reshape(1, -1, img_ids.shape[-1])

        freqs = self.rope_embedder(img_ids).movedim(1, 2)
@@ -643,7 +636,7 @@ class WanModel(torch.nn.Module):
        if self.ref_conv is not None and "reference_latent" in kwargs:
            t_len += 1

-        freqs = self.rope_encode(t_len, h, w, device=x.device, dtype=x.dtype, transformer_options=transformer_options)
+        freqs = self.rope_encode(t_len, h, w, device=x.device, dtype=x.dtype)
        return self.forward_orig(x, timestep, context, clip_fea=clip_fea, freqs=freqs, transformer_options=transformer_options, **kwargs)[:, :, :t, :h, :w]

    def unpatchify(self, x, grid_sizes):
@@ -701,7 +694,7 @@ class VaceWanModel(WanModel):
                 operations=None,
                 ):

-        super().__init__(model_type='t2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, image_model=image_model, device=device, dtype=dtype, operations=operations)
+        super().__init__(model_type='t2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, image_model=image_model, device=device, dtype=dtype, operations=operations, enable_flipflop=False)
        operation_settings = {"operations": operations, "device": device, "dtype": dtype}

        # Vace
@@ -821,7 +814,7 @@ class CameraWanModel(WanModel):
        else:
            model_type = 't2v'

-        super().__init__(model_type=model_type, patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, image_model=image_model, device=device, dtype=dtype, operations=operations)
+        super().__init__(model_type=model_type, patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, image_model=image_model, device=device, dtype=dtype, operations=operations, enable_flipflop=False)
        operation_settings = {"operations": operations, "device": device, "dtype": dtype}

        self.control_adapter = WanCamAdapter(in_dim_control_adapter, dim, kernel_size=patch_size[1:], stride=patch_size[1:], operation_settings=operation_settings)
@@ -1224,7 +1217,7 @@ class WanModel_S2V(WanModel):
                 operations=None,
                 ):

-        super().__init__(model_type='t2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, image_model=image_model, device=device, dtype=dtype, operations=operations)
+        super().__init__(model_type='t2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, image_model=image_model, device=device, dtype=dtype, operations=operations, enable_flipflop=False)

        self.trainable_cond_mask = operations.Embedding(3, self.dim, device=device, dtype=dtype)

@@ -1524,7 +1517,7 @@ class HumoWanModel(WanModel):
                 operations=None,
                 ):

-        super().__init__(model_type='t2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, wan_attn_block_class=WanAttentionBlockAudio, image_model=image_model, device=device, dtype=dtype, operations=operations)
+        super().__init__(model_type='t2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, wan_attn_block_class=WanAttentionBlockAudio, image_model=image_model, device=device, dtype=dtype, operations=operations, enable_flipflop=False)

        self.audio_proj = AudioProjModel(seq_len=8, blocks=5, channels=1280, intermediate_dim=512, output_dim=1536, context_tokens=audio_token_num, dtype=dtype, device=device, operations=operations)

--- a/comfy/ldm/wan/model_animate.py
+++ b/comfy/ldm/wan/model_animate.py
@@ -426,7 +426,7 @@ class AnimateWanModel(WanModel):
                 operations=None,
                 ):

-        super().__init__(model_type='i2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, image_model=image_model, device=device, dtype=dtype, operations=operations)
+        super().__init__(model_type='i2v', patch_size=patch_size, text_len=text_len, in_dim=in_dim, dim=dim, ffn_dim=ffn_dim, freq_dim=freq_dim, text_dim=text_dim, out_dim=out_dim, num_heads=num_heads, num_layers=num_layers, window_size=window_size, qk_norm=qk_norm, cross_attn_norm=cross_attn_norm, eps=eps, flf_pos_embed_token_number=flf_pos_embed_token_number, image_model=image_model, device=device, dtype=dtype, operations=operations, enable_flipflop=False)

        self.pose_patch_embedding = operations.Conv3d(
            16, dim, kernel_size=patch_size, stride=patch_size, device=device, dtype=dtype
--- a/comfy/lora.py
+++ b/comfy/lora.py
@@ -313,15 +313,6 @@ def model_lora_keys_unet(model, key_map={}):
                key_map["transformer.{}".format(key_lora)] = k
                key_map["lycoris_{}".format(key_lora.replace(".", "_"))] = k #SimpleTuner lycoris format

-    if isinstance(model, comfy.model_base.Lumina2):
-        diffusers_keys = comfy.utils.z_image_to_diffusers(model.model_config.unet_config, output_prefix="diffusion_model.")
-        for k in diffusers_keys:
-            if k.endswith(".weight"):
-                to = diffusers_keys[k]
-                key_lora = k[:-len(".weight")]
-                key_map["diffusion_model.{}".format(key_lora)] = to
-                key_map["lycoris_{}".format(key_lora.replace(".", "_"))] = to
-
    return key_map


--- a/comfy/model_base.py
+++ b/comfy/model_base.py
@@ -898,13 +898,12 @@ class Flux(BaseModel):
        attention_mask = kwargs.get("attention_mask", None)
        if attention_mask is not None:
            shape = kwargs["noise"].shape
-            mask_ref_size = kwargs.get("attention_mask_img_shape", None)
-            if mask_ref_size is not None:
-                # the model will pad to the patch size, and then divide
-                # essentially dividing and rounding up
-                (h_tok, w_tok) = (math.ceil(shape[2] / self.diffusion_model.patch_size), math.ceil(shape[3] / self.diffusion_model.patch_size))
-                attention_mask = utils.upscale_dit_mask(attention_mask, mask_ref_size, (h_tok, w_tok))
-                out['attention_mask'] = comfy.conds.CONDRegular(attention_mask)
+            mask_ref_size = kwargs["attention_mask_img_shape"]
+            # the model will pad to the patch size, and then divide
+            # essentially dividing and rounding up
+            (h_tok, w_tok) = (math.ceil(shape[2] / self.diffusion_model.patch_size), math.ceil(shape[3] / self.diffusion_model.patch_size))
+            attention_mask = utils.upscale_dit_mask(attention_mask, mask_ref_size, (h_tok, w_tok))
+            out['attention_mask'] = comfy.conds.CONDRegular(attention_mask)

        guidance = kwargs.get("guidance", 3.5)
        if guidance is not None:
@@ -926,19 +925,9 @@ class Flux(BaseModel):
        out = {}
        ref_latents = kwargs.get("reference_latents", None)
        if ref_latents is not None:
-            out['ref_latents'] = list([1, 16, sum(map(lambda a: math.prod(a.size()[2:]), ref_latents))])
+            out['ref_latents'] = list([1, 16, sum(map(lambda a: math.prod(a.size()), ref_latents)) // 16])
        return out

-class Flux2(Flux):
-    def extra_conds(self, **kwargs):
-        out = super().extra_conds(**kwargs)
-        cross_attn = kwargs.get("cross_attn", None)
-        if cross_attn is not None:
-            target_text_len = 512
-            if cross_attn.shape[1] < target_text_len:
-                cross_attn = torch.nn.functional.pad(cross_attn, (0, 0, target_text_len - cross_attn.shape[1], 0))
-            out['c_crossattn'] = comfy.conds.CONDRegular(cross_attn)
-        return out

 class GenmoMochi(BaseModel):
    def __init__(self, model_config, model_type=ModelType.FLOW, device=None):
@@ -1114,13 +1103,9 @@ class Lumina2(BaseModel):
            if torch.numel(attention_mask) != attention_mask.sum():
                out['attention_mask'] = comfy.conds.CONDRegular(attention_mask)
            out['num_tokens'] = comfy.conds.CONDConstant(max(1, torch.sum(attention_mask).item()))
-
        cross_attn = kwargs.get("cross_attn", None)
        if cross_attn is not None:
            out['c_crossattn'] = comfy.conds.CONDRegular(cross_attn)
-            if 'num_tokens' not in out:
-                out['num_tokens'] = comfy.conds.CONDConstant(cross_attn.shape[1])
-
        return out

 class WAN21(BaseModel):
@@ -1551,94 +1536,3 @@ class HunyuanImage21Refiner(HunyuanImage21):
        out = super().extra_conds(**kwargs)
        out['disable_time_r'] = comfy.conds.CONDConstant(True)
        return out
-
-class HunyuanVideo15(HunyuanVideo):
-    def __init__(self, model_config, model_type=ModelType.FLOW, device=None):
-        super().__init__(model_config, model_type, device=device)
-
-    def concat_cond(self, **kwargs):
-        noise = kwargs.get("noise", None)
-        extra_channels = self.diffusion_model.img_in.proj.weight.shape[1] - noise.shape[1] - 1 #noise 32 img cond 32 + mask 1
-        if extra_channels == 0:
-            return None
-
-        image = kwargs.get("concat_latent_image", None)
-        device = kwargs["device"]
-
-        if image is None:
-            shape_image = list(noise.shape)
-            shape_image[1] = extra_channels
-            image = torch.zeros(shape_image, dtype=noise.dtype, layout=noise.layout, device=noise.device)
-        else:
-            latent_dim = self.latent_format.latent_channels
-            image = utils.common_upscale(image.to(device), noise.shape[-1], noise.shape[-2], "bilinear", "center")
-            for i in range(0, image.shape[1], latent_dim):
-                image[:, i: i + latent_dim] = self.process_latent_in(image[:, i: i + latent_dim])
-            image = utils.resize_to_batch_size(image, noise.shape[0])
-
-        mask = kwargs.get("concat_mask", kwargs.get("denoise_mask", None))
-        if mask is None:
-            mask = torch.zeros_like(noise)[:, :1]
-        else:
-            mask = 1.0 - mask
-            mask = utils.common_upscale(mask.to(device), noise.shape[-1], noise.shape[-2], "bilinear", "center")
-            if mask.shape[-3] < noise.shape[-3]:
-                mask = torch.nn.functional.pad(mask, (0, 0, 0, 0, 0, noise.shape[-3] - mask.shape[-3]), mode='constant', value=0)
-            mask = utils.resize_to_batch_size(mask, noise.shape[0])
-
-        return torch.cat((image, mask), dim=1)
-
-    def extra_conds(self, **kwargs):
-        out = super().extra_conds(**kwargs)
-        attention_mask = kwargs.get("attention_mask", None)
-        if attention_mask is not None:
-            if torch.numel(attention_mask) != attention_mask.sum():
-                out['attention_mask'] = comfy.conds.CONDRegular(attention_mask)
-        cross_attn = kwargs.get("cross_attn", None)
-        if cross_attn is not None:
-            out['c_crossattn'] = comfy.conds.CONDRegular(cross_attn)
-
-        conditioning_byt5small = kwargs.get("conditioning_byt5small", None)
-        if conditioning_byt5small is not None:
-            out['txt_byt5'] = comfy.conds.CONDRegular(conditioning_byt5small)
-
-        guidance = kwargs.get("guidance", 6.0)
-        if guidance is not None:
-            out['guidance'] = comfy.conds.CONDRegular(torch.FloatTensor([guidance]))
-
-        clip_vision_output = kwargs.get("clip_vision_output", None)
-        if clip_vision_output is not None:
-            out['clip_fea'] = comfy.conds.CONDRegular(clip_vision_output.last_hidden_state)
-
-        return out
-
-class HunyuanVideo15_SR_Distilled(HunyuanVideo15):
-    def __init__(self, model_config, model_type=ModelType.FLOW, device=None):
-        super().__init__(model_config, model_type, device=device)
-
-    def concat_cond(self, **kwargs):
-        noise = kwargs.get("noise", None)
-        image = kwargs.get("concat_latent_image", None)
-        noise_augmentation = kwargs.get("noise_augmentation", 0.0)
-        device = kwargs["device"]
-
-        if image is None:
-            image = torch.zeros([noise.shape[0], noise.shape[1] * 2 + 2, noise.shape[-3], noise.shape[-2], noise.shape[-1]], device=comfy.model_management.intermediate_device())
-        else:
-            image = utils.common_upscale(image.to(device), noise.shape[-1], noise.shape[-2], "bilinear", "center")
-            #image = self.process_latent_in(image) # scaling wasn't applied in reference code
-            image = utils.resize_to_batch_size(image, noise.shape[0])
-            lq_image_slice = slice(noise.shape[1] + 1, 2 * noise.shape[1] + 1)
-            if noise_augmentation > 0:
-                generator = torch.Generator(device="cpu")
-                generator.manual_seed(kwargs.get("seed", 0) - 10)
-                noise = torch.randn(image[:, lq_image_slice].shape, generator=generator, dtype=image.dtype, device="cpu").to(image.device)
-                image[:, lq_image_slice] = noise_augmentation * noise + min(1.0 - noise_augmentation, 0.75) * image[:, lq_image_slice]
-            else:
-                image[:, lq_image_slice] = 0.75 * image[:, lq_image_slice]
-        return image
-
-    def extra_conds(self, **kwargs):
-        out = super().extra_conds(**kwargs)
-        out['disable_time_r'] = comfy.conds.CONDConstant(False)
-        return out
--- a/comfy/model_detection.py
+++ b/comfy/model_detection.py
@@ -186,68 +186,30 @@ def detect_unet_config(state_dict, key_prefix, metadata=None):

        guidance_keys = list(filter(lambda a: a.startswith("{}guidance_in.".format(key_prefix)), state_dict_keys))
        dit_config["guidance_embed"] = len(guidance_keys) > 0
-
-        # HunyuanVideo 1.5
-        if '{}cond_type_embedding.weight'.format(key_prefix) in state_dict_keys:
-            dit_config["use_cond_type_embedding"] = True
-        else:
-            dit_config["use_cond_type_embedding"] = False
-        if '{}vision_in.proj.0.weight'.format(key_prefix) in state_dict_keys:
-            dit_config["vision_in_dim"] = state_dict['{}vision_in.proj.0.weight'.format(key_prefix)].shape[0]
-        else:
-            dit_config["vision_in_dim"] = None
        return dit_config

    if '{}double_blocks.0.img_attn.norm.key_norm.scale'.format(key_prefix) in state_dict_keys and ('{}img_in.weight'.format(key_prefix) in state_dict_keys or f"{key_prefix}distilled_guidance_layer.norms.0.scale" in state_dict_keys): #Flux, Chroma or Chroma Radiance (has no img_in.weight)
        dit_config = {}
-        if '{}double_stream_modulation_img.lin.weight'.format(key_prefix) in state_dict_keys:
-            dit_config["image_model"] = "flux2"
-            dit_config["axes_dim"] = [32, 32, 32, 32]
-            dit_config["num_heads"] = 48
-            dit_config["mlp_ratio"] = 3.0
-            dit_config["theta"] = 2000
-            dit_config["out_channels"] = 128
-            dit_config["global_modulation"] = True
-            dit_config["vec_in_dim"] = None
-            dit_config["mlp_silu_act"] = True
-            dit_config["qkv_bias"] = False
-            dit_config["ops_bias"] = False
-            dit_config["default_ref_method"] = "index"
-            dit_config["ref_index_scale"] = 10.0
-            patch_size = 1
-        else:
-            dit_config["image_model"] = "flux"
-            dit_config["axes_dim"] = [16, 56, 56]
-            dit_config["num_heads"] = 24
-            dit_config["mlp_ratio"] = 4.0
-            dit_config["theta"] = 10000
-            dit_config["out_channels"] = 16
-            dit_config["qkv_bias"] = True
-            patch_size = 2
-
+        dit_config["image_model"] = "flux"
        dit_config["in_channels"] = 16
-        dit_config["hidden_size"] = 3072
-        dit_config["context_in_dim"] = 4096
-
+        patch_size = 2
        dit_config["patch_size"] = patch_size
        in_key = "{}img_in.weight".format(key_prefix)
        if in_key in state_dict_keys:
-            w = state_dict[in_key]
-            dit_config["in_channels"] = w.shape[1] // (patch_size * patch_size)
-            dit_config["hidden_size"] = w.shape[0]
-
-        txt_in_key = "{}txt_in.weight".format(key_prefix)
-        if txt_in_key in state_dict_keys:
-            w = state_dict[txt_in_key]
-            dit_config["context_in_dim"] = w.shape[1]
-            dit_config["hidden_size"] = w.shape[0]
-
+            dit_config["in_channels"] = state_dict[in_key].shape[1] // (patch_size * patch_size)
+        dit_config["out_channels"] = 16
        vec_in_key = '{}vector_in.in_layer.weight'.format(key_prefix)
        if vec_in_key in state_dict_keys:
            dit_config["vec_in_dim"] = state_dict[vec_in_key].shape[1]
-
+        dit_config["context_in_dim"] = 4096
+        dit_config["hidden_size"] = 3072
+        dit_config["mlp_ratio"] = 4.0
+        dit_config["num_heads"] = 24
        dit_config["depth"] = count_blocks(state_dict_keys, '{}double_blocks.'.format(key_prefix) + '{}.')
        dit_config["depth_single_blocks"] = count_blocks(state_dict_keys, '{}single_blocks.'.format(key_prefix) + '{}.')
+        dit_config["axes_dim"] = [16, 56, 56]
+        dit_config["theta"] = 10000
+        dit_config["qkv_bias"] = True
        if '{}distilled_guidance_layer.0.norms.0.scale'.format(key_prefix) in state_dict_keys or '{}distilled_guidance_layer.norms.0.scale'.format(key_prefix) in state_dict_keys: #Chroma
            dit_config["image_model"] = "chroma"
            dit_config["in_channels"] = 64
@@ -416,31 +378,14 @@ def detect_unet_config(state_dict, key_prefix, metadata=None):
        dit_config["image_model"] = "lumina2"
        dit_config["patch_size"] = 2
        dit_config["in_channels"] = 16
-        w = state_dict['{}cap_embedder.1.weight'.format(key_prefix)]
-        dit_config["dim"] = w.shape[0]
-        dit_config["cap_feat_dim"] = w.shape[1]
+        dit_config["dim"] = 2304
+        dit_config["cap_feat_dim"] = state_dict['{}cap_embedder.1.weight'.format(key_prefix)].shape[1]
        dit_config["n_layers"] = count_blocks(state_dict_keys, '{}layers.'.format(key_prefix) + '{}.')
+        dit_config["n_heads"] = 24
+        dit_config["n_kv_heads"] = 8
        dit_config["qk_norm"] = True
-
-        if dit_config["dim"] == 2304: # Original Lumina 2
-            dit_config["n_heads"] = 24
-            dit_config["n_kv_heads"] = 8
-            dit_config["axes_dims"] = [32, 32, 32]
-            dit_config["axes_lens"] = [300, 512, 512]
-            dit_config["rope_theta"] = 10000.0
-            dit_config["ffn_dim_multiplier"] = 4.0
-        elif dit_config["dim"] == 3840:  # Z image
-            dit_config["n_heads"] = 30
-            dit_config["n_kv_heads"] = 30
-            dit_config["axes_dims"] = [32, 48, 48]
-            dit_config["axes_lens"] = [1536, 512, 512]
-            dit_config["rope_theta"] = 256.0
-            dit_config["ffn_dim_multiplier"] = (8.0 / 3.0)
-            dit_config["z_image_modulation"] = True
-            dit_config["time_scale"] = 1000.0
-            if '{}cap_pad_token'.format(key_prefix) in state_dict_keys:
-                dit_config["pad_tokens_multiple"] = 32
-
+        dit_config["axes_dims"] = [32, 32, 32]
+        dit_config["axes_lens"] = [300, 512, 512]
        return dit_config

    if '{}head.modulation'.format(key_prefix) in state_dict_keys:  # Wan 2.1
--- a/comfy/model_management.py
+++ b/comfy/model_management.py
@@ -504,7 +504,6 @@ class LoadedModel:
        if use_more_vram == 0:
            use_more_vram = 1e32
        self.model_use_more_vram(use_more_vram, force_patch_weights=force_patch_weights)
-
        real_model = self.model.model

        if is_intel_xpu() and not args.disable_ipex_optimize and 'ipex' in globals() and real_model is not None:
@@ -689,11 +688,8 @@ def load_models_gpu(models, memory_required=0, force_patch_weights=False, minimu
            loaded_memory = loaded_model.model_loaded_memory()
            current_free_mem = get_free_memory(torch_dev) + loaded_memory

-            lowvram_model_memory = max(0, (current_free_mem - minimum_memory_required), min(current_free_mem * MIN_WEIGHT_MEMORY_RATIO, current_free_mem - minimum_inference_memory()))
-            lowvram_model_memory = lowvram_model_memory - loaded_memory
-
-            if lowvram_model_memory == 0:
-                lowvram_model_memory = 0.1
+            lowvram_model_memory = max(128 * 1024 * 1024, (current_free_mem - minimum_memory_required), min(current_free_mem * MIN_WEIGHT_MEMORY_RATIO, current_free_mem - minimum_inference_memory()))
+            lowvram_model_memory = max(0.1, lowvram_model_memory - loaded_memory)

        if vram_set_state == VRAMState.NO_VRAM:
            lowvram_model_memory = 0.1
@@ -1010,74 +1006,58 @@ def force_channels_last():
    #TODO
    return False

+def flipflop_enabled():
+    return args.flipflop_offload

 STREAMS = {}
-NUM_STREAMS = 0
-if args.async_offload is not None:
-    NUM_STREAMS = args.async_offload
-else:
-    #  Enable by default on Nvidia
-    if is_nvidia():
-        NUM_STREAMS = 2
-
-if args.disable_async_offload:
-    NUM_STREAMS = 0
-
-if NUM_STREAMS > 0:
+NUM_STREAMS = 1
+if args.async_offload:
+    NUM_STREAMS = 2
    logging.info("Using async weight offloading with {} streams".format(NUM_STREAMS))

-def current_stream(device):
-    if device is None:
-        return None
-    if is_device_cuda(device):
-        return torch.cuda.current_stream()
-    elif is_device_xpu(device):
-        return torch.xpu.current_stream()
-    else:
-        return None
-
 stream_counters = {}
 def get_offload_stream(device):
    stream_counter = stream_counters.get(device, 0)
-    if NUM_STREAMS == 0:
-        return None
-
-    if torch.compiler.is_compiling():
+    if NUM_STREAMS <= 1:
        return None

    if device in STREAMS:
        ss = STREAMS[device]
-        #Sync the oldest stream in the queue with the current
-        ss[stream_counter].wait_stream(current_stream(device))
+        s = ss[stream_counter]
        stream_counter = (stream_counter + 1) % len(ss)
+        if is_device_cuda(device):
+            ss[stream_counter].wait_stream(torch.cuda.current_stream())
+        elif is_device_xpu(device):
+            ss[stream_counter].wait_stream(torch.xpu.current_stream())
        stream_counters[device] = stream_counter
-        return ss[stream_counter]
+        return s
    elif is_device_cuda(device):
        ss = []
        for k in range(NUM_STREAMS):
-            s1 = torch.cuda.Stream(device=device, priority=0)
-            s1.as_context = torch.cuda.stream
-            ss.append(s1)
+            ss.append(torch.cuda.Stream(device=device, priority=0))
        STREAMS[device] = ss
        s = ss[stream_counter]
+        stream_counter = (stream_counter + 1) % len(ss)
        stream_counters[device] = stream_counter
        return s
    elif is_device_xpu(device):
        ss = []
        for k in range(NUM_STREAMS):
-            s1 = torch.xpu.Stream(device=device, priority=0)
-            s1.as_context = torch.xpu.stream
-            ss.append(s1)
+            ss.append(torch.xpu.Stream(device=device, priority=0))
        STREAMS[device] = ss
        s = ss[stream_counter]
+        stream_counter = (stream_counter + 1) % len(ss)
        stream_counters[device] = stream_counter
        return s
    return None

 def sync_stream(device, stream):
-    if stream is None or current_stream(device) is None:
+    if stream is None:
        return
-    current_stream(device).wait_stream(stream)
+    if is_device_cuda(device):
+        torch.cuda.current_stream().wait_stream(stream)
+    elif is_device_xpu(device):
+        torch.xpu.current_stream().wait_stream(stream)

 def cast_to(weight, dtype=None, device=None, non_blocking=False, copy=False, stream=None):
    if device is None or weight.device == device:
@@ -1085,19 +1065,12 @@ def cast_to(weight, dtype=None, device=None, non_blocking=False, copy=False, str
            if dtype is None or weight.dtype == dtype:
                return weight
        if stream is not None:
-            wf_context = stream
-            if hasattr(wf_context, "as_context"):
-                wf_context = wf_context.as_context(stream)
-            with wf_context:
+            with stream:
                return weight.to(dtype=dtype, copy=copy)
        return weight.to(dtype=dtype, copy=copy)

-
    if stream is not None:
-        wf_context = stream
-        if hasattr(wf_context, "as_context"):
-            wf_context = wf_context.as_context(stream)
-        with wf_context:
+        with stream:
            r = torch.empty_like(weight, dtype=dtype, device=device)
            r.copy_(weight, non_blocking=non_blocking)
    else:
@@ -1109,83 +1082,6 @@ def cast_to_device(tensor, device, dtype, copy=False):
    non_blocking = device_supports_non_blocking(device)
    return cast_to(tensor, dtype=dtype, device=device, non_blocking=non_blocking, copy=copy)

-
-PINNED_MEMORY = {}
-TOTAL_PINNED_MEMORY = 0
-MAX_PINNED_MEMORY = -1
-if not args.disable_pinned_memory:
-    if is_nvidia() or is_amd():
-        if WINDOWS:
-            MAX_PINNED_MEMORY = get_total_memory(torch.device("cpu")) * 0.45  # Windows limit is apparently 50%
-        else:
-            MAX_PINNED_MEMORY = get_total_memory(torch.device("cpu")) * 0.95
-        logging.info("Enabled pinned memory {}".format(MAX_PINNED_MEMORY // (1024 * 1024)))
-
-PINNING_ALLOWED_TYPES = set(["Parameter", "QuantizedTensor"])
-
-def pin_memory(tensor):
-    global TOTAL_PINNED_MEMORY
-    if MAX_PINNED_MEMORY <= 0:
-        return False
-
-    if type(tensor).__name__ not in PINNING_ALLOWED_TYPES:
-        return False
-
-    if not is_device_cpu(tensor.device):
-        return False
-
-    if tensor.is_pinned():
-        #NOTE: Cuda does detect when a tensor is already pinned and would
-        #error below, but there are proven cases where this also queues an error
-        #on the GPU async. So dont trust the CUDA API and guard here
-        return False
-
-    if not tensor.is_contiguous():
-        return False
-
-    size = tensor.numel() * tensor.element_size()
-    if (TOTAL_PINNED_MEMORY + size) > MAX_PINNED_MEMORY:
-        return False
-
-    ptr = tensor.data_ptr()
-    if ptr == 0:
-        return False
-
-    if torch.cuda.cudart().cudaHostRegister(ptr, size, 1) == 0:
-        PINNED_MEMORY[ptr] = size
-        TOTAL_PINNED_MEMORY += size
-        return True
-
-    return False
-
-def unpin_memory(tensor):
-    global TOTAL_PINNED_MEMORY
-    if MAX_PINNED_MEMORY <= 0:
-        return False
-
-    if not is_device_cpu(tensor.device):
-        return False
-
-    ptr = tensor.data_ptr()
-    size = tensor.numel() * tensor.element_size()
-
-    size_stored = PINNED_MEMORY.get(ptr, None)
-    if size_stored is None:
-        logging.warning("Tried to unpin tensor not pinned by ComfyUI")
-        return False
-
-    if size != size_stored:
-        logging.warning("Size of pinned tensor changed")
-        return False
-
-    if torch.cuda.cudart().cudaHostUnregister(ptr) == 0:
-        TOTAL_PINNED_MEMORY -= PINNED_MEMORY.pop(ptr)
-        if len(PINNED_MEMORY) == 0:
-            TOTAL_PINNED_MEMORY = 0
-        return True
-
-    return False
-
 def sage_attention_enabled():
    return args.use_sage_attention

--- a/comfy/model_patcher.py
+++ b/comfy/model_patcher.py
@@ -25,7 +25,7 @@ import logging
 import math
 import uuid
 from typing import Callable, Optional
-
+import time # TODO remove
 import torch

 import comfy.float
@@ -132,7 +132,7 @@ class LowVramPatch:
    def __call__(self, weight):
        intermediate_dtype = weight.dtype
        if self.convert_func is not None:
-            weight = self.convert_func(weight, inplace=False)
+            weight = self.convert_func(weight.to(dtype=torch.float32, copy=True), inplace=True)

        if intermediate_dtype not in [torch.float32, torch.float16, torch.bfloat16]: #intermediate_dtype has to be one that is supported in math ops
            intermediate_dtype = torch.float32
@@ -148,15 +148,6 @@ class LowVramPatch:
        else:
            return out

-#The above patch logic may cast up the weight to fp32, and do math. Go with fp32 x 3
-LOWVRAM_PATCH_ESTIMATE_MATH_FACTOR = 3
-
-def low_vram_patch_estimate_vram(model, key):
-    weight, set_func, convert_func = get_key_weight(model, key)
-    if weight is None:
-        return 0
-    return weight.numel() * torch.float32.itemsize * LOWVRAM_PATCH_ESTIMATE_MATH_FACTOR
-
 def get_key_weight(model, key):
    set_func = None
    convert_func = None
@@ -240,13 +231,13 @@ class ModelPatcher:
        self.object_patches_backup = {}
        self.weight_wrapper_patches = {}
        self.model_options = {"transformer_options":{}}
+        self.model_size()
        self.load_device = load_device
        self.offload_device = offload_device
        self.weight_inplace_update = weight_inplace_update
        self.force_cast_weights = False
        self.patches_uuid = uuid.uuid4()
        self.parent = None
-        self.pinned = set()

        self.attachments: dict[str] = {}
        self.additional_models: dict[str, list[ModelPatcher]] = {}
@@ -278,18 +269,12 @@ class ModelPatcher:
        if not hasattr(self.model, 'current_weight_patches_uuid'):
            self.model.current_weight_patches_uuid = None

-        if not hasattr(self.model, 'model_offload_buffer_memory'):
-            self.model.model_offload_buffer_memory = 0
-
    def model_size(self):
        if self.size > 0:
            return self.size
        self.size = comfy.model_management.module_size(self.model)
        return self.size

-    def get_ram_usage(self):
-        return self.model_size()
-
    def loaded_size(self):
        return self.model.model_loaded_weight_memory

@@ -297,7 +282,7 @@ class ModelPatcher:
        return self.model.lowvram_patch_counter

    def clone(self):
-        n = self.__class__(self.model, self.load_device, self.offload_device, self.model_size(), weight_inplace_update=self.weight_inplace_update)
+        n = self.__class__(self.model, self.load_device, self.offload_device, self.size, weight_inplace_update=self.weight_inplace_update)
        n.patches = {}
        for k in self.patches:
            n.patches[k] = self.patches[k][:]
@@ -309,7 +294,6 @@ class ModelPatcher:
        n.backup = self.backup
        n.object_patches_backup = self.object_patches_backup
        n.parent = self
-        n.pinned = self.pinned

        n.force_cast_weights = self.force_cast_weights

@@ -466,19 +450,6 @@ class ModelPatcher:
    def set_model_post_input_patch(self, patch):
        self.set_model_patch(patch, "post_input")

-    def set_model_rope_options(self, scale_x, shift_x, scale_y, shift_y, scale_t, shift_t, **kwargs):
-        rope_options = self.model_options["transformer_options"].get("rope_options", {})
-        rope_options["scale_x"] = scale_x
-        rope_options["scale_y"] = scale_y
-        rope_options["scale_t"] = scale_t
-
-        rope_options["shift_x"] = shift_x
-        rope_options["shift_y"] = shift_y
-        rope_options["shift_t"] = shift_t
-
-        self.model_options["transformer_options"]["rope_options"] = rope_options
-
-
    def add_object_patch(self, name, obj):
        self.object_patches[name] = obj

@@ -620,7 +591,7 @@ class ModelPatcher:
                        sd.pop(k)
            return sd

-    def patch_weight_to_device(self, key, device_to=None, inplace_update=False):
+    def patch_weight_to_device(self, key, device_to=None, inplace_update=False, device_final=None):
        if key not in self.patches:
            return

@@ -640,30 +611,103 @@ class ModelPatcher:
        out_weight = comfy.lora.calculate_weight(self.patches[key], temp_weight, key)
        if set_func is None:
            out_weight = comfy.float.stochastic_rounding(out_weight, weight.dtype, seed=string_to_seed(key))
+            if device_final is not None:
+                out_weight = out_weight.to(device_final)
            if inplace_update:
                comfy.utils.copy_to_param(self.model, key, out_weight)
            else:
                comfy.utils.set_attr_param(self.model, key, out_weight)
        else:
+            if device_final is not None:
+                out_weight = out_weight.to(device_final)
            set_func(out_weight, inplace_update=inplace_update, seed=string_to_seed(key))

-    def pin_weight_to_device(self, key):
-        weight, set_func, convert_func = get_key_weight(self.model, key)
-        if comfy.model_management.pin_memory(weight):
-            self.pinned.add(key)
+    def supports_flipflop(self):
+        # flipflop requires diffusion_model, explicit flipflop support, NVIDIA CUDA streams, and loading/offloading VRAM
+        if not comfy.model_management.flipflop_enabled():
+            return False
+        if not hasattr(self.model, "diffusion_model"):
+            return False
+        if not getattr(self.model.diffusion_model, "enable_flipflop", False):
+            return False
+        if not comfy.model_management.is_nvidia():
+            return False
+        if comfy.model_management.vram_state in (comfy.model_management.VRAMState.HIGH_VRAM, comfy.model_management.VRAMState.SHARED):
+            return False
+        return True

-    def unpin_weight(self, key):
-        if key in self.pinned:
-            weight, set_func, convert_func = get_key_weight(self.model, key)
-            comfy.model_management.unpin_memory(weight)
-            self.pinned.remove(key)
+    def setup_flipflop(self, flipflop_blocks_per_type: dict[str, tuple[int, int]], flipflop_prefixes: list[str]):
+        if not self.supports_flipflop():
+            return
+        logging.info(f"setting up flipflop with {flipflop_blocks_per_type}")
+        self.model.diffusion_model.setup_flipflop_holders(flipflop_blocks_per_type, flipflop_prefixes, self.load_device, self.offload_device)

-    def unpin_all_weights(self):
-        for key in list(self.pinned):
-            self.unpin_weight(key)
+    def init_flipflop_block_copies(self) -> int:
+        if not self.supports_flipflop():
+            return 0
+        return self.model.diffusion_model.init_flipflop_block_copies(self.load_device)

-    def _load_list(self):
+    def clean_flipflop(self) -> int:
+        if not self.supports_flipflop():
+            return 0
+        return self.model.diffusion_model.clean_flipflop_holders()
+
+    def _get_existing_flipflop_prefixes(self):
+        if self.supports_flipflop():
+            return self.model.diffusion_model.flipflop_prefixes
+        return []
+
+    def _calc_flipflop_prefixes(self, lowvram_model_memory=0, prepare_flipflop=False):
+        flipflop_prefixes = []
+        flipflop_blocks_per_type: dict[str, tuple[int, int]] = {}
+        if lowvram_model_memory > 0 and self.supports_flipflop():
+            block_buffer = 3
+            valid_block_types = []
+            # for each block type, check if have enough room to flipflop
+            for block_info in self.model.diffusion_model.get_all_block_module_sizes(reverse_sort_by_size=True):
+                block_size: int = block_info[1]
+                if block_size * block_buffer < lowvram_model_memory:
+                    valid_block_types.append(block_info)
+            # if have candidates for flipping, see how many of each type we have can flipflop
+            if len(valid_block_types) > 0:
+                leftover_memory = lowvram_model_memory
+                for block_info in valid_block_types:
+                    block_type: str = block_info[0]
+                    block_size: int = block_info[1]
+                    total_blocks = len(self.model.diffusion_model.get_all_blocks(block_type))
+                    n_fit_in_memory = int(leftover_memory // block_size)
+                    # if all (or more) of this block type would fit in memory, no need to flipflop with it
+                    if n_fit_in_memory >= total_blocks:
+                        leftover_memory -= total_blocks * block_size
+                        continue
+                    # if the amount of this block that would fit in memory is less than buffer, skip this block type
+                    if n_fit_in_memory < block_buffer:
+                        continue
+                    # 2 blocks worth of VRAM may be needed for flipflop, so make sure to account for them.
+                    flipflop_blocks = min((total_blocks - n_fit_in_memory) + 2, total_blocks)
+                    # for now, work around odd number issue by making it even
+                    if flipflop_blocks % 2 != 0:
+                        if flipflop_blocks == total_blocks:
+                            flipflop_blocks -= 1
+                        else:
+                            flipflop_blocks += 1
+                    flipflop_blocks_per_type[block_type] = (flipflop_blocks, total_blocks)
+                    leftover_memory -= (total_blocks - flipflop_blocks + 2) * block_size
+                # if there are blocks to flipflop, need to mark their keys
+                for block_type, (flipflop_blocks, total_blocks) in flipflop_blocks_per_type.items():
+                    # blocks to flipflop are at the end
+                    for i in range(total_blocks-flipflop_blocks, total_blocks):
+                        flipflop_prefixes.append(f"diffusion_model.{block_type}.{i}")
+        if prepare_flipflop and len(flipflop_blocks_per_type) > 0:
+            self.setup_flipflop(flipflop_blocks_per_type, flipflop_prefixes)
+        return flipflop_prefixes
+
+    def _load_list(self, lowvram_model_memory=0, prepare_flipflop=False, get_existing_flipflop=False):
        loading = []
+        if get_existing_flipflop:
+            flipflop_prefixes = self._get_existing_flipflop_prefixes()
+        else:
+            flipflop_prefixes = self._calc_flipflop_prefixes(lowvram_model_memory, prepare_flipflop)
        for n, m in self.model.named_modules():
            params = []
            skip = False
@@ -674,16 +718,12 @@ class ModelPatcher:
                    skip = True # skip random weights in non leaf modules
                    break
            if not skip and (hasattr(m, "comfy_cast_weights") or len(params) > 0):
-                module_mem = comfy.model_management.module_size(m)
-                module_offload_mem = module_mem
-                if hasattr(m, "comfy_cast_weights"):
-                    weight_key = "{}.weight".format(n)
-                    bias_key = "{}.bias".format(n)
-                    if weight_key in self.patches:
-                        module_offload_mem += low_vram_patch_estimate_vram(self.model, weight_key)
-                    if bias_key in self.patches:
-                        module_offload_mem += low_vram_patch_estimate_vram(self.model, bias_key)
-                loading.append((module_offload_mem, module_mem, n, m, params))
+                flipflop = False
+                for prefix in flipflop_prefixes:
+                    if n.startswith(prefix):
+                        flipflop = True
+                        break
+                loading.append((comfy.model_management.module_size(m), n, m, params, flipflop))
        return loading

    def load(self, device_to=None, lowvram_model_memory=0, force_patch_weights=False, full_load=False):
@@ -693,26 +733,27 @@ class ModelPatcher:
            patch_counter = 0
            lowvram_counter = 0
            lowvram_mem_counter = 0
-            loading = self._load_list()
+            flipflop_counter = 0
+            flipflop_mem_counter = 0
+            loading = self._load_list(lowvram_model_memory, prepare_flipflop=True)

            load_completely = []
-            offloaded = []
-            offload_buffer = 0
+            load_flipflop = []
            loading.sort(reverse=True)
            for x in loading:
-                module_offload_mem, module_mem, n, m, params = x
+                n = x[1]
+                m = x[2]
+                params = x[3]
+                flipflop: bool = x[4]
+                module_mem = x[0]

                lowvram_weight = False

-                potential_offload = max(offload_buffer, module_offload_mem * (comfy.model_management.NUM_STREAMS + 1))
-                lowvram_fits = mem_counter + module_mem + potential_offload < lowvram_model_memory
-
                weight_key = "{}.weight".format(n)
                bias_key = "{}.bias".format(n)

-                if not full_load and hasattr(m, "comfy_cast_weights"):
-                    if not lowvram_fits:
-                        offload_buffer = potential_offload
+                if not full_load and hasattr(m, "comfy_cast_weights") and not flipflop:
+                    if mem_counter + module_mem >= lowvram_model_memory:
                        lowvram_weight = True
                        lowvram_counter += 1
                        lowvram_mem_counter += module_mem
@@ -741,16 +782,17 @@ class ModelPatcher:
                            patch_counter += 1

                    cast_weight = True
-                    offloaded.append((module_mem, n, m, params))
                else:
                    if hasattr(m, "comfy_cast_weights"):
                        wipe_lowvram_weight(m)

-                    if full_load or lowvram_fits:
+                    if flipflop:
+                        flipflop_counter += 1
+                        flipflop_mem_counter += module_mem
+                        load_flipflop.append((module_mem, n, m, params))
+                    elif full_load or mem_counter + module_mem < lowvram_model_memory:
                        mem_counter += module_mem
                        load_completely.append((module_mem, n, m, params))
-                    else:
-                        offload_buffer = potential_offload

                if cast_weight and hasattr(m, "comfy_cast_weights"):
                    m.prev_comfy_cast_weights = m.comfy_cast_weights
@@ -764,6 +806,7 @@ class ModelPatcher:

                mem_counter += move_weight_functions(m, device_to)

+            # handle load completely
            load_completely.sort(reverse=True)
            for x in load_completely:
                n = x[1]
@@ -774,9 +817,7 @@ class ModelPatcher:
                        continue

                for param in params:
-                    key = "{}.{}".format(n, param)
-                    self.unpin_weight(key)
-                    self.patch_weight_to_device(key, device_to=device_to)
+                    self.patch_weight_to_device("{}.{}".format(n, param), device_to=device_to)

                logging.debug("lowvram: loaded module regularly {} {}".format(n, m))
                m.comfy_patched_weights = True
@@ -784,17 +825,36 @@ class ModelPatcher:
            for x in load_completely:
                x[2].to(device_to)

-            for x in offloaded:
-                n = x[1]
-                params = x[3]
-                for param in params:
-                    self.pin_weight_to_device("{}.{}".format(n, param))
+            # handle flipflop
+            if len(load_flipflop) > 0:
+                start_time = time.perf_counter()
+                load_flipflop.sort(reverse=True)
+                for x in load_flipflop:
+                    n = x[1]
+                    m = x[2]
+                    params = x[3]
+                    if hasattr(m, "comfy_patched_weights"):
+                        if m.comfy_patched_weights == True:
+                            continue
+                    for param in params:
+                        self.patch_weight_to_device("{}.{}".format(n, param), device_to=device_to, device_final=self.offload_device)

-            if lowvram_counter > 0:
-                logging.info("loaded partially; {:.2f} MB usable, {:.2f} MB loaded, {:.2f} MB offloaded, {:.2f} MB buffer reserved, lowvram patches: {}".format(lowvram_model_memory / (1024 * 1024), mem_counter / (1024 * 1024), lowvram_mem_counter / (1024 * 1024), offload_buffer / (1024 * 1024), patch_counter))
+                    logging.debug("lowvram: loaded module for flipflop {} {}".format(n, m))
+                end_time = time.perf_counter()
+                logging.info(f"flipflop load time: {end_time - start_time:.2f} seconds")
+                start_time = time.perf_counter()
+                mem_counter += self.init_flipflop_block_copies()
+                end_time = time.perf_counter()
+                logging.info(f"flipflop block init time: {end_time - start_time:.2f} seconds")
+
+            if lowvram_counter > 0 or flipflop_counter > 0:
+                if flipflop_counter > 0:
+                    logging.info(f"loaded partially; {lowvram_model_memory / (1024 * 1024):.2f} MB usable, {mem_counter / (1024 * 1024):.2f} MB loaded, {flipflop_mem_counter / (1024 * 1024):.2f} MB flipflop, {lowvram_mem_counter / (1024 * 1024):.2f} MB offloaded, lowvram patches: {patch_counter}")
+                else:
+                    logging.info(f"loaded partially; {lowvram_model_memory / (1024 * 1024):.2f} MB usable, {mem_counter / (1024 * 1024):.2f} MB loaded, {lowvram_mem_counter / (1024 * 1024):.2f} MB offloaded, lowvram patches: {patch_counter}")
                self.model.model_lowvram = True
            else:
-                logging.info("loaded completely; {:.2f} MB usable, {:.2f} MB loaded, full load: {}".format(lowvram_model_memory / (1024 * 1024), mem_counter / (1024 * 1024), full_load))
+                logging.info(f"loaded completely; {lowvram_model_memory / (1024 * 1024):.2f} MB usable, {mem_counter / (1024 * 1024):.2f} MB loaded, full load: {full_load}")
                self.model.model_lowvram = False
                if full_load:
                    self.model.to(device_to)
@@ -803,7 +863,6 @@ class ModelPatcher:
            self.model.lowvram_patch_counter += patch_counter
            self.model.device = device_to
            self.model.model_loaded_weight_memory = mem_counter
-            self.model.model_offload_buffer_memory = offload_buffer
            self.model.current_weight_patches_uuid = self.patches_uuid

            for callback in self.get_all_callbacks(CallbacksMP.ON_LOAD):
@@ -832,7 +891,7 @@ class ModelPatcher:
        self.eject_model()
        if unpatch_weights:
            self.unpatch_hooks()
-            self.unpin_all_weights()
+            self.clean_flipflop()
            if self.model.model_lowvram:
                for m in self.model.modules():
                    move_weight_functions(m, device_to)
@@ -857,7 +916,6 @@ class ModelPatcher:
                self.model.to(device_to)
                self.model.device = device_to
            self.model.model_loaded_weight_memory = 0
-            self.model.model_offload_buffer_memory = 0

            for m in self.model.modules():
                if hasattr(m, "comfy_patched_weights"):
@@ -869,22 +927,25 @@ class ModelPatcher:

        self.object_patches_backup.clear()

-    def partially_unload(self, device_to, memory_to_free=0, force_patch_weights=False):
+    def partially_unload(self, device_to, memory_to_free=0):
        with self.use_ejected():
            hooks_unpatched = False
            memory_freed = 0
+            memory_freed += self.clean_flipflop()
            patch_counter = 0
-            unload_list = self._load_list()
+            unload_list = self._load_list(get_existing_flipflop=True)
            unload_list.sort()
-            offload_buffer = self.model.model_offload_buffer_memory
-
            for unload in unload_list:
-                if memory_to_free + offload_buffer - self.model.model_offload_buffer_memory < memory_freed:
+                if memory_to_free < memory_freed:
                    break
-                module_offload_mem, module_mem, n, m, params = unload
-
-                potential_offload = (comfy.model_management.NUM_STREAMS + 1) * module_offload_mem
+                module_mem = unload[0]
+                n = unload[1]
+                m = unload[2]
+                params = unload[3]
+                flipflop: bool = unload[4]

+                if flipflop:
+                    continue
                lowvram_possible = hasattr(m, "comfy_cast_weights")
                if hasattr(m, "comfy_patched_weights") and m.comfy_patched_weights == True:
                    move_weight = True
@@ -914,19 +975,13 @@ class ModelPatcher:
                        module_mem += move_weight_functions(m, device_to)
                        if lowvram_possible:
                            if weight_key in self.patches:
-                                if force_patch_weights:
-                                    self.patch_weight_to_device(weight_key)
-                                else:
-                                    _, set_func, convert_func = get_key_weight(self.model, weight_key)
-                                    m.weight_function.append(LowVramPatch(weight_key, self.patches, convert_func, set_func))
-                                    patch_counter += 1
+                                _, set_func, convert_func = get_key_weight(self.model, weight_key)
+                                m.weight_function.append(LowVramPatch(weight_key, self.patches, convert_func, set_func))
+                                patch_counter += 1
                            if bias_key in self.patches:
-                                if force_patch_weights:
-                                    self.patch_weight_to_device(bias_key)
-                                else:
-                                    _, set_func, convert_func = get_key_weight(self.model, bias_key)
-                                    m.bias_function.append(LowVramPatch(bias_key, self.patches, convert_func, set_func))
-                                    patch_counter += 1
+                                _, set_func, convert_func = get_key_weight(self.model, bias_key)
+                                m.bias_function.append(LowVramPatch(bias_key, self.patches, convert_func, set_func))
+                                patch_counter += 1
                            cast_weight = True

                        if cast_weight:
@@ -934,18 +989,11 @@ class ModelPatcher:
                            m.comfy_cast_weights = True
                        m.comfy_patched_weights = False
                        memory_freed += module_mem
-                        offload_buffer = max(offload_buffer, potential_offload)
                        logging.debug("freed {}".format(n))

-                        for param in params:
-                            self.pin_weight_to_device("{}.{}".format(n, param))
-
-
            self.model.model_lowvram = True
            self.model.lowvram_patch_counter += patch_counter
            self.model.model_loaded_weight_memory -= memory_freed
-            self.model.model_offload_buffer_memory = offload_buffer
-            logging.info("Unloaded partially: {:.2f} MB freed, {:.2f} MB remains loaded, {:.2f} MB buffer reserved, lowvram patches: {}".format(memory_freed / (1024 * 1024), self.model.model_loaded_weight_memory / (1024 * 1024), offload_buffer / (1024 * 1024), self.model.lowvram_patch_counter))
            return memory_freed

    def partially_load(self, device_to, extra_memory=0, force_patch_weights=False):
@@ -958,9 +1006,6 @@ class ModelPatcher:
                extra_memory += (used - self.model.model_loaded_weight_memory)

            self.patch_model(load_weights=False)
-            if extra_memory < 0 and not unpatch_weights:
-                self.partially_unload(self.offload_device, -extra_memory, force_patch_weights=force_patch_weights)
-                return 0
            full_load = False
            if self.model.model_lowvram == False and self.model.model_loaded_weight_memory > 0:
                self.apply_hooks(self.forced_hooks, force_apply=True)
@@ -1348,6 +1393,5 @@ class ModelPatcher:
        self.clear_cached_hook_weights()

    def __del__(self):
-        self.unpin_all_weights()
        self.detach(unpatch_all=False)

--- a/comfy/ops.py
+++ b/comfy/ops.py
@@ -35,7 +35,7 @@ def scaled_dot_product_attention(q, k, v, *args, **kwargs):


 try:
-    if torch.cuda.is_available() and comfy.model_management.WINDOWS:
+    if torch.cuda.is_available():
        from torch.nn.attention import SDPBackend, sdpa_kernel
        import inspect
        if "set_priority" in inspect.signature(sdpa_kernel).parameters:
@@ -58,8 +58,7 @@ except (ModuleNotFoundError, TypeError):
 NVIDIA_MEMORY_CONV_BUG_WORKAROUND = False
 try:
    if comfy.model_management.is_nvidia():
-        cudnn_version = torch.backends.cudnn.version()
-        if (cudnn_version >= 91002 and cudnn_version < 91500) and comfy.model_management.torch_version_numeric >= (2, 9) and comfy.model_management.torch_version_numeric <= (2, 10):
+        if torch.backends.cudnn.version() >= 91002 and comfy.model_management.torch_version_numeric >= (2, 9) and comfy.model_management.torch_version_numeric <= (2, 10):
            #TODO: change upper bound version once it's fixed'
            NVIDIA_MEMORY_CONV_BUG_WORKAROUND = True
            logging.info("working around nvidia conv3d memory bug.")
@@ -71,78 +70,42 @@ cast_to = comfy.model_management.cast_to #TODO: remove once no more references
 def cast_to_input(weight, input, non_blocking=False, copy=True):
    return comfy.model_management.cast_to(weight, input.dtype, input.device, non_blocking=non_blocking, copy=copy)

-
-def cast_bias_weight(s, input=None, dtype=None, device=None, bias_dtype=None, offloadable=False):
-    # NOTE: offloadable=False is a a legacy and if you are a custom node author reading this please pass
-    # offloadable=True and call uncast_bias_weight() after your last usage of the weight/bias. This
-    # will add async-offload support to your cast and improve performance.
+@torch.compiler.disable()
+def cast_bias_weight(s, input=None, dtype=None, device=None, bias_dtype=None):
    if input is not None:
        if dtype is None:
-            if isinstance(input, QuantizedTensor):
-                dtype = input._layout_params["orig_dtype"]
-            else:
-                dtype = input.dtype
+            dtype = input.dtype
        if bias_dtype is None:
            bias_dtype = dtype
        if device is None:
            device = input.device

-    if offloadable and (device != s.weight.device or
-                        (s.bias is not None and device != s.bias.device)):
-        offload_stream = comfy.model_management.get_offload_stream(device)
-    else:
-        offload_stream = None
-
+    offload_stream = comfy.model_management.get_offload_stream(device)
    if offload_stream is not None:
        wf_context = offload_stream
-        if hasattr(wf_context, "as_context"):
-            wf_context = wf_context.as_context(offload_stream)
    else:
        wf_context = contextlib.nullcontext()

-    non_blocking = comfy.model_management.device_supports_non_blocking(device)
-
-    weight_has_function = len(s.weight_function) > 0
-    bias_has_function = len(s.bias_function) > 0
-
-    weight = comfy.model_management.cast_to(s.weight, None, device, non_blocking=non_blocking, copy=weight_has_function, stream=offload_stream)
-
    bias = None
+    non_blocking = comfy.model_management.device_supports_non_blocking(device)
    if s.bias is not None:
-        bias = comfy.model_management.cast_to(s.bias, bias_dtype, device, non_blocking=non_blocking, copy=bias_has_function, stream=offload_stream)
+        has_function = len(s.bias_function) > 0
+        bias = comfy.model_management.cast_to(s.bias, bias_dtype, device, non_blocking=non_blocking, copy=has_function, stream=offload_stream)

-        if bias_has_function:
+        if has_function:
            with wf_context:
                for f in s.bias_function:
                    bias = f(bias)

-    if weight_has_function or weight.dtype != dtype:
+    has_function = len(s.weight_function) > 0
+    weight = comfy.model_management.cast_to(s.weight, dtype, device, non_blocking=non_blocking, copy=has_function, stream=offload_stream)
+    if has_function:
        with wf_context:
-            weight = weight.to(dtype=dtype)
-            if isinstance(weight, QuantizedTensor):
-                weight = weight.dequantize()
            for f in s.weight_function:
                weight = f(weight)

    comfy.model_management.sync_stream(device, offload_stream)
-    if offloadable:
-        return weight, bias, offload_stream
-    else:
-        #Legacy function signature
-        return weight, bias
-
-
-def uncast_bias_weight(s, weight, bias, offload_stream):
-    if offload_stream is None:
-        return
-    if weight is not None:
-        device = weight.device
-    else:
-        if bias is None:
-            return
-        device = bias.device
-    offload_stream.wait_stream(comfy.model_management.current_stream(device))
-
+    return weight, bias

 class CastWeightBiasOp:
    comfy_cast_weights = False
@@ -155,10 +118,8 @@ class disable_weight_init:
            return None

        def forward_comfy_cast_weights(self, input):
-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = torch.nn.functional.linear(input, weight, bias)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            weight, bias = cast_bias_weight(self, input)
+            return torch.nn.functional.linear(input, weight, bias)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -172,10 +133,8 @@ class disable_weight_init:
            return None

        def forward_comfy_cast_weights(self, input):
-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = self._conv_forward(input, weight, bias)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            weight, bias = cast_bias_weight(self, input)
+            return self._conv_forward(input, weight, bias)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -189,10 +148,8 @@ class disable_weight_init:
            return None

        def forward_comfy_cast_weights(self, input):
-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = self._conv_forward(input, weight, bias)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            weight, bias = cast_bias_weight(self, input)
+            return self._conv_forward(input, weight, bias)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -215,10 +172,8 @@ class disable_weight_init:
                return super()._conv_forward(input, weight, bias, *args, **kwargs)

        def forward_comfy_cast_weights(self, input):
-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = self._conv_forward(input, weight, bias)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            weight, bias = cast_bias_weight(self, input)
+            return self._conv_forward(input, weight, bias)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -232,10 +187,8 @@ class disable_weight_init:
            return None

        def forward_comfy_cast_weights(self, input):
-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = torch.nn.functional.group_norm(input, self.num_groups, weight, bias, self.eps)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            weight, bias = cast_bias_weight(self, input)
+            return torch.nn.functional.group_norm(input, self.num_groups, weight, bias, self.eps)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -250,14 +203,11 @@ class disable_weight_init:

        def forward_comfy_cast_weights(self, input):
            if self.weight is not None:
-                weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
+                weight, bias = cast_bias_weight(self, input)
            else:
                weight = None
                bias = None
-                offload_stream = None
-            x = torch.nn.functional.layer_norm(input, self.normalized_shape, weight, bias, self.eps)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            return torch.nn.functional.layer_norm(input, self.normalized_shape, weight, bias, self.eps)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -273,15 +223,11 @@ class disable_weight_init:

        def forward_comfy_cast_weights(self, input):
            if self.weight is not None:
-                weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
+                weight, bias = cast_bias_weight(self, input)
            else:
                weight = None
-                bias = None
-                offload_stream = None
-            x = comfy.rmsnorm.rms_norm(input, weight, self.eps)  # TODO: switch to commented out line when old torch is deprecated
-            # x = torch.nn.functional.rms_norm(input, self.normalized_shape, weight, self.eps)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            return comfy.rmsnorm.rms_norm(input, weight, self.eps)  # TODO: switch to commented out line when old torch is deprecated
+            # return torch.nn.functional.rms_norm(input, self.normalized_shape, weight, self.eps)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -300,12 +246,10 @@ class disable_weight_init:
                input, output_size, self.stride, self.padding, self.kernel_size,
                num_spatial_dims, self.dilation)

-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = torch.nn.functional.conv_transpose2d(
+            weight, bias = cast_bias_weight(self, input)
+            return torch.nn.functional.conv_transpose2d(
                input, weight, bias, self.stride, self.padding,
                output_padding, self.groups, self.dilation)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -324,12 +268,10 @@ class disable_weight_init:
                input, output_size, self.stride, self.padding, self.kernel_size,
                num_spatial_dims, self.dilation)

-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = torch.nn.functional.conv_transpose1d(
+            weight, bias = cast_bias_weight(self, input)
+            return torch.nn.functional.conv_transpose1d(
                input, weight, bias, self.stride, self.padding,
                output_padding, self.groups, self.dilation)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -347,11 +289,8 @@ class disable_weight_init:
            output_dtype = out_dtype
            if self.weight.dtype == torch.float16 or self.weight.dtype == torch.bfloat16:
                out_dtype = None
-            weight, bias, offload_stream = cast_bias_weight(self, device=input.device, dtype=out_dtype, offloadable=True)
-            x = torch.nn.functional.embedding(input, weight, self.padding_idx, self.max_norm, self.norm_type, self.scale_grad_by_freq, self.sparse).to(dtype=output_dtype)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
-
+            weight, bias = cast_bias_weight(self, device=input.device, dtype=out_dtype)
+            return torch.nn.functional.embedding(input, weight, self.padding_idx, self.max_norm, self.norm_type, self.scale_grad_by_freq, self.sparse).to(dtype=output_dtype)

        def forward(self, *args, **kwargs):
            run_every_op()
@@ -413,10 +352,16 @@ def fp8_linear(self, input):
    if dtype not in [torch.float8_e4m3fn]:
        return None

+    tensor_2d = False
+    if len(input.shape) == 2:
+        tensor_2d = True
+        input = input.unsqueeze(1)
+
+    input_shape = input.shape
    input_dtype = input.dtype

-    if input.ndim == 3 or input.ndim == 2:
-        w, bias, offload_stream = cast_bias_weight(self, input, dtype=dtype, bias_dtype=input_dtype, offloadable=True)
+    if len(input.shape) == 3:
+        w, bias = cast_bias_weight(self, input, dtype=dtype, bias_dtype=input_dtype)

        scale_weight = self.scale_weight
        scale_input = self.scale_input
@@ -427,21 +372,19 @@ def fp8_linear(self, input):

        if scale_input is None:
            scale_input = torch.ones((), device=input.device, dtype=torch.float32)
-            input = torch.clamp(input, min=-448, max=448, out=input)
-            layout_params_weight = {'scale': scale_input, 'orig_dtype': input_dtype}
-            quantized_input = QuantizedTensor(input.to(dtype).contiguous(), "TensorCoreFP8Layout", layout_params_weight)
        else:
            scale_input = scale_input.to(input.device)
-            quantized_input = QuantizedTensor.from_float(input, "TensorCoreFP8Layout", scale=scale_input, dtype=dtype)

        # Wrap weight in QuantizedTensor - this enables unified dispatch
        # Call F.linear - __torch_dispatch__ routes to fp8_linear handler in quant_ops.py!
        layout_params_weight = {'scale': scale_weight, 'orig_dtype': input_dtype}
-        quantized_weight = QuantizedTensor(w, "TensorCoreFP8Layout", layout_params_weight)
+        quantized_weight = QuantizedTensor(w, TensorCoreFP8Layout, layout_params_weight)
+        quantized_input = QuantizedTensor.from_float(input.reshape(-1, input_shape[2]), TensorCoreFP8Layout, scale=scale_input, dtype=dtype)
        o = torch.nn.functional.linear(quantized_input, quantized_weight, bias)

-        uncast_bias_weight(self, w, bias, offload_stream)
-        return o
+        if tensor_2d:
+            return o.reshape(input_shape[0], -1)
+        return o.reshape((-1, input_shape[1], self.weight.shape[0]))

    return None

@@ -461,10 +404,8 @@ class fp8_ops(manual_cast):
                except Exception as e:
                    logging.info("Exception during fp8 op: {}".format(e))

-            weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-            x = torch.nn.functional.linear(input, weight, bias)
-            uncast_bias_weight(self, weight, bias, offload_stream)
-            return x
+            weight, bias = cast_bias_weight(self, input)
+            return torch.nn.functional.linear(input, weight, bias)

 def scaled_fp8_ops(fp8_matrix_mult=False, scale_input=False, override_dtype=None):
    logging.info("Using scaled fp8: fp8 matrix mult: {}, scale input: {}".format(fp8_matrix_mult, scale_input))
@@ -492,21 +433,19 @@ def scaled_fp8_ops(fp8_matrix_mult=False, scale_input=False, override_dtype=None
                    if out is not None:
                        return out

-                weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
+                weight, bias = cast_bias_weight(self, input)

                if weight.numel() < input.numel(): #TODO: optimize
-                    x = torch.nn.functional.linear(input, weight * self.scale_weight.to(device=weight.device, dtype=weight.dtype), bias)
+                    return torch.nn.functional.linear(input, weight * self.scale_weight.to(device=weight.device, dtype=weight.dtype), bias)
                else:
-                    x = torch.nn.functional.linear(input * self.scale_weight.to(device=weight.device, dtype=weight.dtype), weight, bias)
-                uncast_bias_weight(self, weight, bias, offload_stream)
-                return x
+                    return torch.nn.functional.linear(input * self.scale_weight.to(device=weight.device, dtype=weight.dtype), weight, bias)

            def convert_weight(self, weight, inplace=False, **kwargs):
                if inplace:
                    weight *= self.scale_weight.to(device=weight.device, dtype=weight.dtype)
                    return weight
                else:
-                    return weight.to(dtype=torch.float32) * self.scale_weight.to(device=weight.device, dtype=torch.float32)
+                    return weight * self.scale_weight.to(device=weight.device, dtype=weight.dtype)

            def set_weight(self, weight, inplace_update=False, seed=None, return_weight=False, **kwargs):
                weight = comfy.float.stochastic_rounding(weight / self.scale_weight.to(device=weight.device, dtype=weight.dtype), self.weight.dtype, seed=seed)
@@ -542,138 +481,125 @@ if CUBLAS_IS_AVAILABLE:
 # ==============================================================================
 # Mixed Precision Operations
 # ==============================================================================
-from .quant_ops import QuantizedTensor, QUANT_ALGOS
+from .quant_ops import QuantizedTensor, TensorCoreFP8Layout

+QUANT_FORMAT_MIXINS = {
+    "float8_e4m3fn": {
+        "dtype": torch.float8_e4m3fn,
+        "layout_type": TensorCoreFP8Layout,
+        "parameters": {
+            "weight_scale": torch.nn.Parameter(torch.zeros((), dtype=torch.float32), requires_grad=False),
+            "input_scale": torch.nn.Parameter(torch.zeros((), dtype=torch.float32), requires_grad=False),
+        }
+    }
+}

-def mixed_precision_ops(layer_quant_config={}, compute_dtype=torch.bfloat16, full_precision_mm=False):
-    class MixedPrecisionOps(manual_cast):
-        _layer_quant_config = layer_quant_config
-        _compute_dtype = compute_dtype
-        _full_precision_mm = full_precision_mm
+class MixedPrecisionOps(disable_weight_init):
+    _layer_quant_config = {}
+    _compute_dtype = torch.bfloat16

-        class Linear(torch.nn.Module, CastWeightBiasOp):
-            def __init__(
-                self,
-                in_features: int,
-                out_features: int,
-                bias: bool = True,
-                device=None,
-                dtype=None,
-            ) -> None:
-                super().__init__()
+    class Linear(torch.nn.Module, CastWeightBiasOp):
+        def __init__(
+            self,
+            in_features: int,
+            out_features: int,
+            bias: bool = True,
+            device=None,
+            dtype=None,
+        ) -> None:
+            super().__init__()

-                self.factory_kwargs = {"device": device, "dtype": MixedPrecisionOps._compute_dtype}
-                # self.factory_kwargs = {"device": device, "dtype": dtype}
+            self.factory_kwargs = {"device": device, "dtype": MixedPrecisionOps._compute_dtype}
+            # self.factory_kwargs = {"device": device, "dtype": dtype}

-                self.in_features = in_features
-                self.out_features = out_features
-                if bias:
-                    self.bias = torch.nn.Parameter(torch.empty(out_features, **self.factory_kwargs))
-                else:
-                    self.register_parameter("bias", None)
+            self.in_features = in_features
+            self.out_features = out_features
+            if bias:
+                self.bias = torch.nn.Parameter(torch.empty(out_features, **self.factory_kwargs))
+            else:
+                self.register_parameter("bias", None)

-                self.tensor_class = None
-                self._full_precision_mm = MixedPrecisionOps._full_precision_mm
+            self.tensor_class = None

-            def reset_parameters(self):
-                return None
+        def reset_parameters(self):
+            return None

-            def _load_from_state_dict(self, state_dict, prefix, local_metadata,
-                                    strict, missing_keys, unexpected_keys, error_msgs):
+        def _load_from_state_dict(self, state_dict, prefix, local_metadata,
+                                  strict, missing_keys, unexpected_keys, error_msgs):

-                device = self.factory_kwargs["device"]
-                layer_name = prefix.rstrip('.')
-                weight_key = f"{prefix}weight"
-                weight = state_dict.pop(weight_key, None)
-                if weight is None:
-                    raise ValueError(f"Missing weight for layer {layer_name}")
+            device = self.factory_kwargs["device"]
+            layer_name = prefix.rstrip('.')
+            weight_key = f"{prefix}weight"
+            weight = state_dict.pop(weight_key, None)
+            if weight is None:
+                raise ValueError(f"Missing weight for layer {layer_name}")

-                manually_loaded_keys = [weight_key]
+            manually_loaded_keys = [weight_key]

-                if layer_name not in MixedPrecisionOps._layer_quant_config:
-                    self.weight = torch.nn.Parameter(weight.to(device=device, dtype=MixedPrecisionOps._compute_dtype), requires_grad=False)
-                else:
-                    quant_format = MixedPrecisionOps._layer_quant_config[layer_name].get("format", None)
-                    if quant_format is None:
-                        raise ValueError(f"Unknown quantization format for layer {layer_name}")
+            if layer_name not in MixedPrecisionOps._layer_quant_config:
+                self.weight = torch.nn.Parameter(weight.to(device=device, dtype=MixedPrecisionOps._compute_dtype), requires_grad=False)
+            else:
+                quant_format = MixedPrecisionOps._layer_quant_config[layer_name].get("format", None)
+                if quant_format is None:
+                    raise ValueError(f"Unknown quantization format for layer {layer_name}")

-                    qconfig = QUANT_ALGOS[quant_format]
-                    self.layout_type = qconfig["comfy_tensor_layout"]
+                mixin = QUANT_FORMAT_MIXINS[quant_format]
+                self.layout_type = mixin["layout_type"]

-                    weight_scale_key = f"{prefix}weight_scale"
-                    layout_params = {
-                        'scale': state_dict.pop(weight_scale_key, None),
-                        'orig_dtype': MixedPrecisionOps._compute_dtype,
-                        'block_size': qconfig.get("group_size", None),
-                    }
-                    if layout_params['scale'] is not None:
-                        manually_loaded_keys.append(weight_scale_key)
+                scale_key = f"{prefix}weight_scale"
+                layout_params = {
+                    'scale': state_dict.pop(scale_key, None),
+                    'orig_dtype': MixedPrecisionOps._compute_dtype
+                }
+                if layout_params['scale'] is not None:
+                    manually_loaded_keys.append(scale_key)

-                    self.weight = torch.nn.Parameter(
-                        QuantizedTensor(weight.to(device=device), self.layout_type, layout_params),
-                        requires_grad=False
-                    )
+                self.weight = torch.nn.Parameter(
+                    QuantizedTensor(weight.to(device=device, dtype=mixin["dtype"]), self.layout_type, layout_params),
+                    requires_grad=False
+                )

-                    for param_name in qconfig["parameters"]:
-                        param_key = f"{prefix}{param_name}"
-                        _v = state_dict.pop(param_key, None)
-                        if _v is None:
-                            continue
-                        setattr(self, param_name, torch.nn.Parameter(_v.to(device=device), requires_grad=False))
-                        manually_loaded_keys.append(param_key)
+                for param_name, param_value in mixin["parameters"].items():
+                    param_key = f"{prefix}{param_name}"
+                    _v = state_dict.pop(param_key, None)
+                    if _v is None:
+                        continue
+                    setattr(self, param_name, torch.nn.Parameter(_v.to(device=device), requires_grad=False))
+                    manually_loaded_keys.append(param_key)

-                super()._load_from_state_dict(state_dict, prefix, local_metadata, strict, missing_keys, unexpected_keys, error_msgs)
+            super()._load_from_state_dict(state_dict, prefix, local_metadata, strict, missing_keys, unexpected_keys, error_msgs)

-                for key in manually_loaded_keys:
-                    if key in missing_keys:
-                        missing_keys.remove(key)
+            for key in manually_loaded_keys:
+                if key in missing_keys:
+                    missing_keys.remove(key)

-            def _forward(self, input, weight, bias):
-                return torch.nn.functional.linear(input, weight, bias)
+        def _forward(self, input, weight, bias):
+            return torch.nn.functional.linear(input, weight, bias)

-            def forward_comfy_cast_weights(self, input):
-                weight, bias, offload_stream = cast_bias_weight(self, input, offloadable=True)
-                x = self._forward(input, weight, bias)
-                uncast_bias_weight(self, weight, bias, offload_stream)
-                return x
+        def forward_comfy_cast_weights(self, input):
+            weight, bias = cast_bias_weight(self, input)
+            return self._forward(input, weight, bias)

-            def forward(self, input, *args, **kwargs):
-                run_every_op()
+        def forward(self, input, *args, **kwargs):
+            run_every_op()

-                if self._full_precision_mm or self.comfy_cast_weights or len(self.weight_function) > 0 or len(self.bias_function) > 0:
-                    return self.forward_comfy_cast_weights(input, *args, **kwargs)
-                if (getattr(self, 'layout_type', None) is not None and
-                    getattr(self, 'input_scale', None) is not None and
-                    not isinstance(input, QuantizedTensor)):
-                    input = QuantizedTensor.from_float(input, self.layout_type, scale=self.input_scale, dtype=self.weight.dtype)
-                return self._forward(input, self.weight, self.bias)
+            if self.comfy_cast_weights or len(self.weight_function) > 0 or len(self.bias_function) > 0:
+                return self.forward_comfy_cast_weights(input, *args, **kwargs)
+            if (getattr(self, 'layout_type', None) is not None and
+                getattr(self, 'input_scale', None) is not None and
+                not isinstance(input, QuantizedTensor)):
+                input = QuantizedTensor.from_float(input, self.layout_type, scale=self.input_scale, fp8_dtype=self.weight.dtype)
+            return self._forward(input, self.weight, self.bias)

-            def convert_weight(self, weight, inplace=False, **kwargs):
-                if isinstance(weight, QuantizedTensor):
-                    return weight.dequantize()
-                else:
-                    return weight
-
-            def set_weight(self, weight, inplace_update=False, seed=None, return_weight=False, **kwargs):
-                if getattr(self, 'layout_type', None) is not None:
-                    weight = QuantizedTensor.from_float(weight, self.layout_type, scale=None, dtype=self.weight.dtype, stochastic_rounding=seed, inplace_ops=True)
-                else:
-                    weight = weight.to(self.weight.dtype)
-                if return_weight:
-                    return weight
-
-                assert inplace_update is False  # TODO: eventually remove the inplace_update stuff
-                self.weight = torch.nn.Parameter(weight, requires_grad=False)
-
-    return MixedPrecisionOps

 def pick_operations(weight_dtype, compute_dtype, load_device=None, disable_fast_fp8=False, fp8_optimizations=False, scaled_fp8=None, model_config=None):
-    fp8_compute = comfy.model_management.supports_fp8_compute(load_device) # TODO: if we support more ops this needs to be more granular
-
    if model_config and hasattr(model_config, 'layer_quant_config') and model_config.layer_quant_config:
+        MixedPrecisionOps._layer_quant_config = model_config.layer_quant_config
+        MixedPrecisionOps._compute_dtype = compute_dtype
        logging.info(f"Using mixed precision operations: {len(model_config.layer_quant_config)} quantized layers")
-        return mixed_precision_ops(model_config.layer_quant_config, compute_dtype, full_precision_mm=not fp8_compute)
+        return MixedPrecisionOps

+    fp8_compute = comfy.model_management.supports_fp8_compute(load_device)
    if scaled_fp8 is not None:
        return scaled_fp8_ops(fp8_matrix_mult=fp8_compute and fp8_optimizations, scale_input=fp8_optimizations, override_dtype=scaled_fp8)

--- a/comfy/quant_ops.py
+++ b/comfy/quant_ops.py
@@ -1,7 +1,6 @@
 import torch
 import logging
 from typing import Tuple, Dict
-import comfy.float

 _LAYOUT_REGISTRY = {}
 _GENERIC_UTILS = {}
@@ -75,12 +74,6 @@ def _copy_layout_params(params):
            new_params[k] = v
    return new_params

-def _copy_layout_params_inplace(src, dst, non_blocking=False):
-    for k, v in src.items():
-        if isinstance(v, torch.Tensor):
-            dst[k].copy_(v, non_blocking=non_blocking)
-        else:
-            dst[k] = v

 class QuantizedLayout:
    """
@@ -130,15 +123,15 @@ class QuantizedTensor(torch.Tensor):
            layout_type: Layout class (subclass of QuantizedLayout)
            layout_params: Dict with layout-specific parameters
        """
-        return torch.Tensor._make_wrapper_subclass(cls, qdata.shape, device=qdata.device, dtype=qdata.dtype, requires_grad=False)
+        return torch.Tensor._make_subclass(cls, qdata, require_grad=False)

    def __init__(self, qdata, layout_type, layout_params):
-        self._qdata = qdata
+        self._qdata = qdata.contiguous()
        self._layout_type = layout_type
        self._layout_params = layout_params

    def __repr__(self):
-        layout_name = self._layout_type
+        layout_name = self._layout_type.__name__
        param_str = ", ".join(f"{k}={v}" for k, v in list(self._layout_params.items())[:2])
        return f"QuantizedTensor(shape={self.shape}, layout={layout_name}, {param_str})"

@@ -186,15 +179,15 @@ class QuantizedTensor(torch.Tensor):
            attr_name = f"_layout_param_{key}"
            layout_params[key] = inner_tensors[attr_name]

-        return QuantizedTensor(inner_tensors["_qdata"], layout_type, layout_params)
+        return QuantizedTensor(inner_tensors["_q_data"], layout_type, layout_params)

    @classmethod
    def from_float(cls, tensor, layout_type, **quantize_kwargs) -> 'QuantizedTensor':
-        qdata, layout_params = LAYOUTS[layout_type].quantize(tensor, **quantize_kwargs)
+        qdata, layout_params = layout_type.quantize(tensor, **quantize_kwargs)
        return cls(qdata, layout_type, layout_params)

    def dequantize(self) -> torch.Tensor:
-        return LAYOUTS[self._layout_type].dequantize(self._qdata, **self._layout_params)
+        return self._layout_type.dequantize(self._qdata, **self._layout_params)

    @classmethod
    def __torch_dispatch__(cls, func, types, args=(), kwargs=None):
@@ -229,14 +222,6 @@ class QuantizedTensor(torch.Tensor):
        new_kwargs = dequant_arg(kwargs)
        return func(*new_args, **new_kwargs)

-    def data_ptr(self):
-        return self._qdata.data_ptr()
-
-    def is_pinned(self):
-        return self._qdata.is_pinned()
-
-    def is_contiguous(self, *arg, **kwargs):
-        return self._qdata.is_contiguous(*arg, **kwargs)

 # ==============================================================================
 # Generic Utilities (Layout-Agnostic Operations)
@@ -333,13 +318,13 @@ def generic_to_dtype_layout(func, args, kwargs):
 def generic_copy_(func, args, kwargs):
    qt_dest = args[0]
    src = args[1]
-    non_blocking = args[2] if len(args) > 2 else False
+
    if isinstance(qt_dest, QuantizedTensor):
        if isinstance(src, QuantizedTensor):
            # Copy from another quantized tensor
-            qt_dest._qdata.copy_(src._qdata, non_blocking=non_blocking)
+            qt_dest._qdata.copy_(src._qdata)
            qt_dest._layout_type = src._layout_type
-            _copy_layout_params_inplace(src._layout_params, qt_dest._layout_params, non_blocking=non_blocking)
+            qt_dest._layout_params = _copy_layout_params(src._layout_params)
        else:
            # Copy from regular tensor - just copy raw data
            qt_dest._qdata.copy_(src)
@@ -347,42 +332,10 @@ def generic_copy_(func, args, kwargs):
    return func(*args, **kwargs)


-@register_generic_util(torch.ops.aten.to.dtype)
-def generic_to_dtype(func, args, kwargs):
-    """Handle .to(dtype) calls - dtype conversion only."""
-    src = args[0]
-    if isinstance(src, QuantizedTensor):
-        # For dtype-only conversion, just change the orig_dtype, no real cast is needed
-        target_dtype = args[1] if len(args) > 1 else kwargs.get('dtype')
-        src._layout_params["orig_dtype"] = target_dtype
-        return src
-    return func(*args, **kwargs)
-
-
@register_generic_util(torch.ops.aten._has_compatible_shallow_copy_type.default)
 def generic_has_compatible_shallow_copy_type(func, args, kwargs):
    return True

-
-@register_generic_util(torch.ops.aten.empty_like.default)
-def generic_empty_like(func, args, kwargs):
-    """Empty_like operation - creates an empty tensor with the same quantized structure."""
-    qt = args[0]
-    if isinstance(qt, QuantizedTensor):
-        # Create empty tensor with same shape and dtype as the quantized data
-        hp_dtype = kwargs.pop('dtype', qt._layout_params["orig_dtype"])
-        new_qdata = torch.empty_like(qt._qdata, **kwargs)
-
-        # Handle device transfer for layout params
-        target_device = kwargs.get('device', new_qdata.device)
-        new_params = _move_layout_params_to_device(qt._layout_params, target_device)
-
-        # Update orig_dtype if dtype is specified
-        new_params['orig_dtype'] = hp_dtype
-
-        return QuantizedTensor(new_qdata, qt._layout_type, new_params)
-    return func(*args, **kwargs)
-
 # ==============================================================================
 # FP8 Layout + Operation Handlers
 # ==============================================================================
@@ -394,7 +347,7 @@ class TensorCoreFP8Layout(QuantizedLayout):
    - orig_dtype: Original dtype before quantization (for casting back)
    """
    @classmethod
-    def quantize(cls, tensor, scale=None, dtype=torch.float8_e4m3fn, stochastic_rounding=0, inplace_ops=False):
+    def quantize(cls, tensor, scale=None, dtype=torch.float8_e4m3fn):
        orig_dtype = tensor.dtype

        if scale is None:
@@ -404,48 +357,28 @@ class TensorCoreFP8Layout(QuantizedLayout):
            scale = torch.tensor(scale)
        scale = scale.to(device=tensor.device, dtype=torch.float32)

-        if inplace_ops:
-            tensor *= (1.0 / scale).to(tensor.dtype)
-        else:
-            tensor = tensor * (1.0 / scale).to(tensor.dtype)
-
-        if stochastic_rounding > 0:
-            tensor = comfy.float.stochastic_rounding(tensor, dtype=dtype, seed=stochastic_rounding)
-        else:
-            lp_amax = torch.finfo(dtype).max
-            torch.clamp(tensor, min=-lp_amax, max=lp_amax, out=tensor)
-            tensor = tensor.to(dtype, memory_format=torch.contiguous_format)
+        lp_amax = torch.finfo(dtype).max
+        tensor_scaled = tensor.float() / scale
+        torch.clamp(tensor_scaled, min=-lp_amax, max=lp_amax, out=tensor_scaled)
+        qdata = tensor_scaled.to(dtype, memory_format=torch.contiguous_format)

        layout_params = {
            'scale': scale,
            'orig_dtype': orig_dtype
        }
-        return tensor, layout_params
+        return qdata, layout_params

    @staticmethod
    def dequantize(qdata, scale, orig_dtype, **kwargs):
        plain_tensor = torch.ops.aten._to_copy.default(qdata, dtype=orig_dtype)
-        plain_tensor.mul_(scale)
-        return plain_tensor
+        return plain_tensor * scale

    @classmethod
    def get_plain_tensors(cls, qtensor):
        return qtensor._qdata, qtensor._layout_params['scale']

-QUANT_ALGOS = {
-    "float8_e4m3fn": {
-        "storage_t": torch.float8_e4m3fn,
-        "parameters": {"weight_scale", "input_scale"},
-        "comfy_tensor_layout": "TensorCoreFP8Layout",
-    },
-}

-LAYOUTS = {
-    "TensorCoreFP8Layout": TensorCoreFP8Layout,
-}
-
-
-@register_layout_op(torch.ops.aten.linear.default, "TensorCoreFP8Layout")
+@register_layout_op(torch.ops.aten.linear.default, TensorCoreFP8Layout)
 def fp8_linear(func, args, kwargs):
    input_tensor = args[0]
    weight = args[1]
@@ -472,17 +405,13 @@ def fp8_linear(func, args, kwargs):

        try:
            output = torch._scaled_mm(
-                plain_input.reshape(-1, input_shape[2]).contiguous(),
+                plain_input.reshape(-1, input_shape[2]),
                weight_t,
                bias=bias,
                scale_a=scale_a,
                scale_b=scale_b,
                out_dtype=out_dtype,
            )
-
-            if isinstance(output, tuple):  # TODO: remove when we drop support for torch 2.4
-                output = output[0]
-
            if not tensor_2d:
                output = output.reshape((-1, input_shape[1], weight.shape[0]))

@@ -492,7 +421,7 @@ def fp8_linear(func, args, kwargs):
                    'scale': output_scale,
                    'orig_dtype': input_tensor._layout_params['orig_dtype']
                }
-                return QuantizedTensor(output, "TensorCoreFP8Layout", output_params)
+                return QuantizedTensor(output, TensorCoreFP8Layout, output_params)
            else:
                return output

@@ -506,68 +435,3 @@ def fp8_linear(func, args, kwargs):
        input_tensor = input_tensor.dequantize()

    return torch.nn.functional.linear(input_tensor, weight, bias)
-
-def fp8_mm_(input_tensor, weight, bias=None, out_dtype=None):
-    if out_dtype is None:
-        out_dtype = input_tensor._layout_params['orig_dtype']
-
-    plain_input, scale_a = TensorCoreFP8Layout.get_plain_tensors(input_tensor)
-    plain_weight, scale_b = TensorCoreFP8Layout.get_plain_tensors(weight)
-
-    output = torch._scaled_mm(
-        plain_input.contiguous(),
-        plain_weight,
-        bias=bias,
-        scale_a=scale_a,
-        scale_b=scale_b,
-        out_dtype=out_dtype,
-    )
-
-    if isinstance(output, tuple):  # TODO: remove when we drop support for torch 2.4
-        output = output[0]
-    return output
-
-@register_layout_op(torch.ops.aten.addmm.default, "TensorCoreFP8Layout")
-def fp8_addmm(func, args, kwargs):
-    input_tensor = args[1]
-    weight = args[2]
-    bias = args[0]
-
-    if isinstance(input_tensor, QuantizedTensor) and isinstance(weight, QuantizedTensor):
-        return fp8_mm_(input_tensor, weight, bias=bias, out_dtype=kwargs.get("out_dtype", None))
-
-    a = list(args)
-    if isinstance(args[0], QuantizedTensor):
-        a[0] = args[0].dequantize()
-    if isinstance(args[1], QuantizedTensor):
-        a[1] = args[1].dequantize()
-    if isinstance(args[2], QuantizedTensor):
-        a[2] = args[2].dequantize()
-
-    return func(*a, **kwargs)
-
-@register_layout_op(torch.ops.aten.mm.default, "TensorCoreFP8Layout")
-def fp8_mm(func, args, kwargs):
-    input_tensor = args[0]
-    weight = args[1]
-
-    if isinstance(input_tensor, QuantizedTensor) and isinstance(weight, QuantizedTensor):
-        return fp8_mm_(input_tensor, weight, bias=None, out_dtype=kwargs.get("out_dtype", None))
-
-    a = list(args)
-    if isinstance(args[0], QuantizedTensor):
-        a[0] = args[0].dequantize()
-    if isinstance(args[1], QuantizedTensor):
-        a[1] = args[1].dequantize()
-    return func(*a, **kwargs)
-
-@register_layout_op(torch.ops.aten.view.default, "TensorCoreFP8Layout")
-@register_layout_op(torch.ops.aten.t.default, "TensorCoreFP8Layout")
-def fp8_func(func, args, kwargs):
-    input_tensor = args[0]
-    if isinstance(input_tensor, QuantizedTensor):
-        plain_input, scale_a = TensorCoreFP8Layout.get_plain_tensors(input_tensor)
-        ar = list(args)
-        ar[0] = plain_input
-        return QuantizedTensor(func(*ar, **kwargs), "TensorCoreFP8Layout", input_tensor._layout_params)
-    return func(*args, **kwargs)
--- a/comfy/sd.py
+++ b/comfy/sd.py
@@ -52,7 +52,6 @@ import comfy.text_encoders.ace
 import comfy.text_encoders.omnigen2
 import comfy.text_encoders.qwen_image
 import comfy.text_encoders.hunyuan_image
-import comfy.text_encoders.z_image

 import comfy.model_patcher
 import comfy.lora
@@ -60,8 +59,6 @@ import comfy.lora_convert
 import comfy.hooks
 import comfy.t2i_adapter.adapter
 import comfy.taesd.taesd
-import comfy.taesd.taehv
-import comfy.latent_formats

 import comfy.ldm.flux.redux

@@ -146,9 +143,6 @@ class CLIP:
        n.apply_hooks_to_conds = self.apply_hooks_to_conds
        return n

-    def get_ram_usage(self):
-        return self.patcher.get_ram_usage()
-
    def add_patches(self, patches, strength_patch=1.0, strength_model=1.0):
        return self.patcher.add_patches(patches, strength_patch, strength_model)

@@ -299,7 +293,6 @@ class VAE:
        self.working_dtypes = [torch.bfloat16, torch.float32]
        self.disable_offload = False
        self.not_video = False
-        self.size = None

        self.downscale_index_formula = None
        self.upscale_index_formula = None
@@ -359,7 +352,7 @@ class VAE:

                    self.memory_used_encode = lambda shape, dtype: (700 * shape[2] * shape[3]) * model_management.dtype_size(dtype)
                    self.memory_used_decode = lambda shape, dtype: (700 * shape[2] * shape[3] * 32 * 32) * model_management.dtype_size(dtype)
-                elif sd['decoder.conv_in.weight'].shape[1] == 32 and sd['decoder.conv_in.weight'].ndim == 5:
+                elif sd['decoder.conv_in.weight'].shape[1] == 32:
                    ddconfig = {"block_out_channels": [128, 256, 512, 1024, 1024], "in_channels": 3, "out_channels": 3, "num_res_blocks": 2, "ffactor_spatial": 16, "ffactor_temporal": 4, "downsample_match_channel": True, "upsample_match_channel": True, "refiner_vae": False}
                    self.latent_channels = ddconfig['z_channels'] = sd["decoder.conv_in.weight"].shape[1]
                    self.working_dtypes = [torch.float16, torch.bfloat16, torch.float32]
@@ -385,17 +378,6 @@ class VAE:
                        self.upscale_ratio = 4

                    self.latent_channels = ddconfig['z_channels'] = sd["decoder.conv_in.weight"].shape[1]
-                    if 'decoder.post_quant_conv.weight' in sd:
-                        sd = comfy.utils.state_dict_prefix_replace(sd, {"decoder.post_quant_conv.": "post_quant_conv.", "encoder.quant_conv.": "quant_conv."})
-
-                    if 'bn.running_mean' in sd:
-                        ddconfig["batch_norm_latent"] = True
-                        self.downscale_ratio *= 2
-                        self.upscale_ratio *= 2
-                        self.latent_channels *= 4
-                        old_memory_used_decode = self.memory_used_decode
-                        self.memory_used_decode = lambda shape, dtype: old_memory_used_decode(shape, dtype) *  4.0
-
                    if 'post_quant_conv.weight' in sd:
                        self.first_stage_model = AutoencoderKL(ddconfig=ddconfig, embed_dim=sd['post_quant_conv.weight'].shape[1])
                    else:
@@ -455,20 +437,20 @@ class VAE:
            elif "decoder.conv_in.conv.weight" in sd and sd['decoder.conv_in.conv.weight'].shape[1] == 32:
                ddconfig = {"block_out_channels": [128, 256, 512, 1024, 1024], "in_channels": 3, "out_channels": 3, "num_res_blocks": 2, "ffactor_spatial": 16, "ffactor_temporal": 4, "downsample_match_channel": True, "upsample_match_channel": True}
                ddconfig['z_channels'] = sd["decoder.conv_in.conv.weight"].shape[1]
-                self.latent_channels = 32
+                self.latent_channels = 64
                self.upscale_ratio = (lambda a: max(0, a * 4 - 3), 16, 16)
                self.upscale_index_formula = (4, 16, 16)
                self.downscale_ratio = (lambda a: max(0, math.floor((a + 3) / 4)), 16, 16)
                self.downscale_index_formula = (4, 16, 16)
                self.latent_dim = 3
-                self.not_video = False
+                self.not_video = True
                self.working_dtypes = [torch.float16, torch.bfloat16, torch.float32]
                self.first_stage_model = AutoencodingEngine(regularizer_config={'target': "comfy.ldm.models.autoencoder.EmptyRegularizer"},
                                                            encoder_config={'target': "comfy.ldm.hunyuan_video.vae_refiner.Encoder", 'params': ddconfig},
                                                            decoder_config={'target': "comfy.ldm.hunyuan_video.vae_refiner.Decoder", 'params': ddconfig})

-                self.memory_used_encode = lambda shape, dtype: (1400 * 9 * shape[-2] * shape[-1]) * model_management.dtype_size(dtype)
-                self.memory_used_decode = lambda shape, dtype: (2800 * 4 * shape[-2] * shape[-1] * 16 * 16) * model_management.dtype_size(dtype)
+                self.memory_used_encode = lambda shape, dtype: (1400 * shape[-2] * shape[-1]) * model_management.dtype_size(dtype)
+                self.memory_used_decode = lambda shape, dtype: (1400 * shape[-3] * shape[-2] * shape[-1] * 16 * 16) * model_management.dtype_size(dtype)
            elif "decoder.conv_in.conv.weight" in sd:
                ddconfig = {'double_z': True, 'z_channels': 4, 'resolution': 256, 'in_channels': 3, 'out_ch': 3, 'ch': 128, 'ch_mult': [1, 2, 4, 4], 'num_res_blocks': 2, 'attn_resolutions': [], 'dropout': 0.0}
                ddconfig["conv3d"] = True
@@ -510,14 +492,13 @@ class VAE:
                    self.memory_used_encode = lambda shape, dtype: 3300 * shape[3] * shape[4] * model_management.dtype_size(dtype)
                    self.memory_used_decode = lambda shape, dtype: 8000 * shape[3] * shape[4] * (16 * 16) * model_management.dtype_size(dtype)
                else:  # Wan 2.1 VAE
-                    dim = sd["decoder.head.0.gamma"].shape[0]
                    self.upscale_ratio = (lambda a: max(0, a * 4 - 3), 8, 8)
                    self.upscale_index_formula = (4, 8, 8)
                    self.downscale_ratio = (lambda a: max(0, math.floor((a + 3) / 4)), 8, 8)
                    self.downscale_index_formula = (4, 8, 8)
                    self.latent_dim = 3
                    self.latent_channels = 16
-                    ddconfig = {"dim": dim, "z_dim": self.latent_channels, "dim_mult": [1, 2, 4, 4], "num_res_blocks": 2, "attn_scales": [], "temperal_downsample": [False, True, True], "dropout": 0.0}
+                    ddconfig = {"dim": 96, "z_dim": self.latent_channels, "dim_mult": [1, 2, 4, 4], "num_res_blocks": 2, "attn_scales": [], "temperal_downsample": [False, True, True], "dropout": 0.0}
                    self.first_stage_model = comfy.ldm.wan.vae.WanVAE(**ddconfig)
                    self.working_dtypes = [torch.bfloat16, torch.float16, torch.float32]
                    self.memory_used_encode = lambda shape, dtype: 6000 * shape[3] * shape[4] * model_management.dtype_size(dtype)
@@ -587,35 +568,6 @@ class VAE:
                self.process_input = lambda audio: audio
                self.working_dtypes = [torch.float32]
                self.crop_input = False
-            elif "decoder.22.bias" in sd: # taehv, taew and lighttae
-                self.latent_channels = sd["decoder.1.weight"].shape[1]
-                self.latent_dim = 3
-                self.upscale_ratio = (lambda a: max(0, a * 4 - 3), 16, 16)
-                self.upscale_index_formula = (4, 16, 16)
-                self.downscale_ratio = (lambda a: max(0, math.floor((a + 3) / 4)), 16, 16)
-                self.downscale_index_formula = (4, 16, 16)
-                if self.latent_channels == 48: # Wan 2.2
-                    self.first_stage_model = comfy.taesd.taehv.TAEHV(latent_channels=self.latent_channels, latent_format=None) # taehv doesn't need scaling
-                    self.process_input = lambda image: (_ for _ in ()).throw(NotImplementedError("This light tae doesn't support encoding currently"))
-                    self.process_output = lambda image: image
-                    self.memory_used_decode = lambda shape, dtype: (1800 * (max(1, (shape[-3] ** 0.7 * 0.1)) * shape[-2] * shape[-1] * 16 * 16) * model_management.dtype_size(dtype))
-                elif self.latent_channels == 32 and sd["decoder.22.bias"].shape[0] == 12: # lighttae_hv15
-                    self.first_stage_model = comfy.taesd.taehv.TAEHV(latent_channels=self.latent_channels, latent_format=comfy.latent_formats.HunyuanVideo15)
-                    self.process_input = lambda image: (_ for _ in ()).throw(NotImplementedError("This light tae doesn't support encoding currently"))
-                    self.memory_used_decode = lambda shape, dtype: (1200 * (max(1, (shape[-3] ** 0.7 * 0.05)) * shape[-2] * shape[-1] * 32 * 32) * model_management.dtype_size(dtype))
-                else:
-                    if sd["decoder.1.weight"].dtype == torch.float16: # taehv currently only available in float16, so assume it's not lighttaew2_1 as otherwise state dicts are identical
-                        latent_format=comfy.latent_formats.HunyuanVideo
-                    else:
-                        latent_format=None # lighttaew2_1 doesn't need scaling
-                    self.first_stage_model = comfy.taesd.taehv.TAEHV(latent_channels=self.latent_channels, latent_format=latent_format)
-                    self.process_input = self.process_output = lambda image: image
-                    self.upscale_ratio = (lambda a: max(0, a * 4 - 3), 8, 8)
-                    self.upscale_index_formula = (4, 8, 8)
-                    self.downscale_ratio = (lambda a: max(0, math.floor((a + 3) / 4)), 8, 8)
-                    self.downscale_index_formula = (4, 8, 8)
-                    self.memory_used_encode = lambda shape, dtype: (700 * (max(1, (shape[-3] ** 0.66 * 0.11)) * shape[-2] * shape[-1]) * model_management.dtype_size(dtype))
-                    self.memory_used_decode = lambda shape, dtype: (50 * (max(1, (shape[-3] ** 0.65 * 0.26)) * shape[-2] * shape[-1] * 32 * 32) * model_management.dtype_size(dtype))
            else:
                logging.warning("WARNING: No VAE weights detected, VAE not initalized.")
                self.first_stage_model = None
@@ -643,16 +595,6 @@ class VAE:

        self.patcher = comfy.model_patcher.ModelPatcher(self.first_stage_model, load_device=self.device, offload_device=offload_device)
        logging.info("VAE load device: {}, offload device: {}, dtype: {}".format(self.device, offload_device, self.vae_dtype))
-        self.model_size()
-
-    def model_size(self):
-        if self.size is not None:
-            return self.size
-        self.size = comfy.model_management.module_size(self.first_stage_model)
-        return self.size
-
-    def get_ram_usage(self):
-        return self.model_size()

    def throw_exception_if_invalid(self):
        if self.first_stage_model is None:
@@ -955,18 +897,12 @@ class CLIPType(Enum):
    OMNIGEN2 = 17
    QWEN_IMAGE = 18
    HUNYUAN_IMAGE = 19
-    HUNYUAN_VIDEO_15 = 20


 def load_clip(ckpt_paths, embedding_directory=None, clip_type=CLIPType.STABLE_DIFFUSION, model_options={}):
    clip_data = []
    for p in ckpt_paths:
-        sd, metadata = comfy.utils.load_torch_file(p, safe_load=True, return_metadata=True)
-        if metadata is not None:
-            quant_metadata = metadata.get("_quantization_metadata", None)
-            if quant_metadata is not None:
-                sd["_quantization_metadata"] = quant_metadata
-        clip_data.append(sd)
+        clip_data.append(comfy.utils.load_torch_file(p, safe_load=True))
    return load_text_encoder_state_dicts(clip_data, embedding_directory=embedding_directory, clip_type=clip_type, model_options=model_options)


@@ -984,10 +920,6 @@ class TEModel(Enum):
    QWEN25_7B = 11
    BYT5_SMALL_GLYPH = 12
    GEMMA_3_4B = 13
-    MISTRAL3_24B = 14
-    MISTRAL3_24B_PRUNED_FLUX2 = 15
-    QWEN3_4B = 16
-

 def detect_te_model(sd):
    if "text_model.encoder.layers.30.mlp.fc1.weight" in sd:
@@ -1020,15 +952,6 @@ def detect_te_model(sd):
        if weight.shape[0] == 512:
            return TEModel.QWEN25_7B
    if "model.layers.0.post_attention_layernorm.weight" in sd:
-        if 'model.layers.0.self_attn.q_norm.weight' in sd:
-            return TEModel.QWEN3_4B
-        weight = sd['model.layers.0.post_attention_layernorm.weight']
-        if weight.shape[0] == 5120:
-            if "model.layers.39.post_attention_layernorm.weight" in sd:
-                return TEModel.MISTRAL3_24B
-            else:
-                return TEModel.MISTRAL3_24B_PRUNED_FLUX2
-
        return TEModel.LLAMA3_8
    return None

@@ -1143,13 +1066,6 @@ def load_text_encoder_state_dicts(state_dicts=[], embedding_directory=None, clip
            else:
                clip_target.clip = comfy.text_encoders.qwen_image.te(**llama_detect(clip_data))
                clip_target.tokenizer = comfy.text_encoders.qwen_image.QwenImageTokenizer
-        elif te_model == TEModel.MISTRAL3_24B or te_model == TEModel.MISTRAL3_24B_PRUNED_FLUX2:
-            clip_target.clip = comfy.text_encoders.flux.flux2_te(**llama_detect(clip_data), pruned=te_model == TEModel.MISTRAL3_24B_PRUNED_FLUX2)
-            clip_target.tokenizer = comfy.text_encoders.flux.Flux2Tokenizer
-            tokenizer_data["tekken_model"] = clip_data[0].get("tekken_model", None)
-        elif te_model == TEModel.QWEN3_4B:
-            clip_target.clip = comfy.text_encoders.z_image.te(**llama_detect(clip_data))
-            clip_target.tokenizer = comfy.text_encoders.z_image.ZImageTokenizer
        else:
            # clip_l
            if clip_type == CLIPType.SD3:
@@ -1196,9 +1112,6 @@ def load_text_encoder_state_dicts(state_dicts=[], embedding_directory=None, clip
        elif clip_type == CLIPType.HUNYUAN_IMAGE:
            clip_target.clip = comfy.text_encoders.hunyuan_image.te(**llama_detect(clip_data))
            clip_target.tokenizer = comfy.text_encoders.hunyuan_image.HunyuanImageTokenizer
-        elif clip_type == CLIPType.HUNYUAN_VIDEO_15:
-            clip_target.clip = comfy.text_encoders.hunyuan_image.te(**llama_detect(clip_data))
-            clip_target.tokenizer = comfy.text_encoders.hunyuan_video.HunyuanVideo15Tokenizer
        else:
            clip_target.clip = sdxl_clip.SDXLClipModel
            clip_target.tokenizer = sdxl_clip.SDXLTokenizer
@@ -1211,8 +1124,6 @@ def load_text_encoder_state_dicts(state_dicts=[], embedding_directory=None, clip

    parameters = 0
    for c in clip_data:
-        if "_quantization_metadata" in c:
-            c.pop("_quantization_metadata")
        parameters += comfy.utils.calculate_parameters(c)
        tokenizer_data, model_options = comfy.text_encoders.long_clipl.model_options_long_clip(c, tokenizer_data, model_options)

@@ -1419,7 +1330,7 @@ def load_diffusion_model_state_dict(sd, model_options={}, metadata=None):
    else:
        unet_dtype = dtype

-    if model_config.layer_quant_config is not None:
+    if hasattr(model_config, "layer_quant_config"):
        manual_cast_dtype = model_management.unet_manual_cast(None, load_device, model_config.supported_inference_dtypes)
    else:
        manual_cast_dtype = model_management.unet_manual_cast(unet_dtype, load_device, model_config.supported_inference_dtypes)
--- a/comfy/sd1_clip.py
+++ b/comfy/sd1_clip.py
@@ -90,6 +90,7 @@ class SDClipModel(torch.nn.Module, ClipTokenWeightEncoder):
                 special_tokens={"start": 49406, "end": 49407, "pad": 49407}, layer_norm_hidden_state=True, enable_attention_masks=False, zero_out_masked=False,
                 return_projected_pooled=True, return_attention_masks=False, model_options={}):  # clip-vit-base-patch32
        super().__init__()
+        assert layer in self.LAYERS

        if textmodel_json_config is None:
            textmodel_json_config = os.path.join(os.path.dirname(os.path.realpath(__file__)), "sd1_clip_config.json")
@@ -108,23 +109,13 @@ class SDClipModel(torch.nn.Module, ClipTokenWeightEncoder):

        operations = model_options.get("custom_operations", None)
        scaled_fp8 = None
-        quantization_metadata = model_options.get("quantization_metadata", None)

        if operations is None:
-            layer_quant_config = None
-            if quantization_metadata is not None:
-                layer_quant_config = json.loads(quantization_metadata).get("layers", None)
-
-            if layer_quant_config is not None:
-                operations = comfy.ops.mixed_precision_ops(layer_quant_config, dtype, full_precision_mm=True)
-                logging.info(f"Using MixedPrecisionOps for text encoder: {len(layer_quant_config)} quantized layers")
+            scaled_fp8 = model_options.get("scaled_fp8", None)
+            if scaled_fp8 is not None:
+                operations = comfy.ops.scaled_fp8_ops(fp8_matrix_mult=False, override_dtype=scaled_fp8)
            else:
-                # Fallback to scaled_fp8_ops for backward compatibility
-                scaled_fp8 = model_options.get("scaled_fp8", None)
-                if scaled_fp8 is not None:
-                    operations = comfy.ops.scaled_fp8_ops(fp8_matrix_mult=False, override_dtype=scaled_fp8)
-                else:
-                    operations = comfy.ops.manual_cast
+                operations = comfy.ops.manual_cast

        self.operations = operations
        self.transformer = model_class(config, dtype, device, self.operations)
@@ -163,7 +154,7 @@ class SDClipModel(torch.nn.Module, ClipTokenWeightEncoder):
    def set_clip_options(self, options):
        layer_idx = options.get("layer", self.layer_idx)
        self.return_projected_pooled = options.get("projected_pooled", self.return_projected_pooled)
-        if isinstance(self.layer, list) or self.layer == "all":
+        if self.layer == "all":
            pass
        elif layer_idx is None or abs(layer_idx) > self.num_layers:
            self.layer = "last"
@@ -265,9 +256,7 @@ class SDClipModel(torch.nn.Module, ClipTokenWeightEncoder):
        if self.enable_attention_masks:
            attention_mask_model = attention_mask

-        if isinstance(self.layer, list):
-            intermediate_output = self.layer
-        elif self.layer == "all":
+        if self.layer == "all":
            intermediate_output = "all"
        else:
            intermediate_output = self.layer_idx
@@ -471,7 +460,7 @@ def load_embed(embedding_name, embedding_directory, embedding_size, embed_key=No
    return embed_out

 class SDTokenizer:
-    def __init__(self, tokenizer_path=None, max_length=77, pad_with_end=True, embedding_directory=None, embedding_size=768, embedding_key='clip_l', tokenizer_class=CLIPTokenizer, has_start_token=True, has_end_token=True, pad_to_max_length=True, min_length=None, pad_token=None, end_token=None, min_padding=None, pad_left=False, tokenizer_data={}, tokenizer_args={}):
+    def __init__(self, tokenizer_path=None, max_length=77, pad_with_end=True, embedding_directory=None, embedding_size=768, embedding_key='clip_l', tokenizer_class=CLIPTokenizer, has_start_token=True, has_end_token=True, pad_to_max_length=True, min_length=None, pad_token=None, end_token=None, min_padding=None, tokenizer_data={}, tokenizer_args={}):
        if tokenizer_path is None:
            tokenizer_path = os.path.join(os.path.dirname(os.path.realpath(__file__)), "sd1_tokenizer")
        self.tokenizer = tokenizer_class.from_pretrained(tokenizer_path, **tokenizer_args)
@@ -479,7 +468,6 @@ class SDTokenizer:
        self.min_length = tokenizer_data.get("{}_min_length".format(embedding_key), min_length)
        self.end_token = None
        self.min_padding = min_padding
-        self.pad_left = pad_left

        empty = self.tokenizer('')["input_ids"]
        self.tokenizer_adds_end_token = has_end_token
@@ -534,12 +522,6 @@ class SDTokenizer:
                return (embed, "{} {}".format(embedding_name[len(stripped):], leftover))
        return (embed, leftover)

-    def pad_tokens(self, tokens, amount):
-        if self.pad_left:
-            for i in range(amount):
-                tokens.insert(0, (self.pad_token, 1.0, 0))
-        else:
-            tokens.extend([(self.pad_token, 1.0, 0)] * amount)

    def tokenize_with_weights(self, text:str, return_word_ids=False, tokenizer_options={}, **kwargs):
        '''
@@ -618,7 +600,7 @@ class SDTokenizer:
                        if self.end_token is not None:
                            batch.append((self.end_token, 1.0, 0))
                        if self.pad_to_max_length:
-                            self.pad_tokens(batch, remaining_length)
+                            batch.extend([(self.pad_token, 1.0, 0)] * (remaining_length))
                    #start new batch
                    batch = []
                    if self.start_token is not None:
@@ -632,11 +614,11 @@ class SDTokenizer:
        if self.end_token is not None:
            batch.append((self.end_token, 1.0, 0))
        if min_padding is not None:
-            self.pad_tokens(batch, min_padding)
+            batch.extend([(self.pad_token, 1.0, 0)] * min_padding)
        if self.pad_to_max_length and len(batch) < self.max_length:
-            self.pad_tokens(batch, self.max_length - len(batch))
+            batch.extend([(self.pad_token, 1.0, 0)] * (self.max_length - len(batch)))
        if min_length is not None and len(batch) < min_length:
-            self.pad_tokens(batch, min_length - len(batch))
+            batch.extend([(self.pad_token, 1.0, 0)] * (min_length - len(batch)))

        if not return_word_ids:
            batched_tokens = [[(t, w) for t, w,_ in x] for x in batched_tokens]
--- a/comfy/supported_models.py
+++ b/comfy/supported_models.py
@@ -21,7 +21,6 @@ import comfy.text_encoders.ace
 import comfy.text_encoders.omnigen2
 import comfy.text_encoders.qwen_image
 import comfy.text_encoders.hunyuan_image
-import comfy.text_encoders.z_image

 from . import supported_models_base
 from . import latent_formats
@@ -742,37 +741,6 @@ class FluxSchnell(Flux):
        out = model_base.Flux(self, model_type=model_base.ModelType.FLOW, device=device)
        return out

-class Flux2(Flux):
-    unet_config = {
-        "image_model": "flux2",
-    }
-
-    sampling_settings = {
-        "shift": 2.02,
-    }
-
-    unet_extra_config = {}
-    latent_format = latent_formats.Flux2
-
-    supported_inference_dtypes = [torch.bfloat16, torch.float16, torch.float32]
-
-    vae_key_prefix = ["vae."]
-    text_encoder_key_prefix = ["text_encoders."]
-
-    def __init__(self, unet_config):
-        super().__init__(unet_config)
-        self.memory_usage_factor = self.memory_usage_factor * (2.0 * 2.0) * 2.36
-
-    def get_model(self, state_dict, prefix="", device=None):
-        out = model_base.Flux2(self, device=device)
-        return out
-
-    def clip_target(self, state_dict={}):
-        return None # TODO
-        pref = self.text_encoder_key_prefix[0]
-        t5_detect = comfy.text_encoders.sd3_clip.t5_xxl_detect(state_dict, "{}t5xxl.transformer.".format(pref))
-        return supported_models_base.ClipTarget(comfy.text_encoders.flux.FluxTokenizer, comfy.text_encoders.flux.flux_clip(**t5_detect))
-
 class GenmoMochi(supported_models_base.BASE):
    unet_config = {
        "image_model": "mochi_preview",
@@ -995,7 +963,7 @@ class Lumina2(supported_models_base.BASE):
        "shift": 6.0,
    }

-    memory_usage_factor = 1.4
+    memory_usage_factor = 1.2

    unet_extra_config = {}
    latent_format = latent_formats.Flux
@@ -1014,24 +982,6 @@ class Lumina2(supported_models_base.BASE):
        hunyuan_detect = comfy.text_encoders.hunyuan_video.llama_detect(state_dict, "{}gemma2_2b.transformer.".format(pref))
        return supported_models_base.ClipTarget(comfy.text_encoders.lumina2.LuminaTokenizer, comfy.text_encoders.lumina2.te(**hunyuan_detect))

-class ZImage(Lumina2):
-    unet_config = {
-        "image_model": "lumina2",
-        "dim": 3840,
-    }
-
-    sampling_settings = {
-        "multiplier": 1.0,
-        "shift": 3.0,
-    }
-
-    memory_usage_factor = 1.7
-
-    def clip_target(self, state_dict={}):
-        pref = self.text_encoder_key_prefix[0]
-        hunyuan_detect = comfy.text_encoders.hunyuan_video.llama_detect(state_dict, "{}qwen3_4b.transformer.".format(pref))
-        return supported_models_base.ClipTarget(comfy.text_encoders.z_image.ZImageTokenizer, comfy.text_encoders.z_image.te(**hunyuan_detect))
-
 class WAN21_T2V(supported_models_base.BASE):
    unet_config = {
        "image_model": "wan2.1",
@@ -1424,55 +1374,6 @@ class HunyuanImage21Refiner(HunyuanVideo):
        out = model_base.HunyuanImage21Refiner(self, device=device)
        return out

-class HunyuanVideo15(HunyuanVideo):
-    unet_config = {
-        "image_model": "hunyuan_video",
-        "vision_in_dim": 1152,
-    }
-
-    sampling_settings = {
-        "shift": 7.0,
-    }
-    memory_usage_factor = 4.0 #TODO
-    supported_inference_dtypes = [torch.float16, torch.bfloat16, torch.float32]
-
-    latent_format = latent_formats.HunyuanVideo15
-
-    def get_model(self, state_dict, prefix="", device=None):
-        out = model_base.HunyuanVideo15(self, device=device)
-        return out
-
-    def clip_target(self, state_dict={}):
-        pref = self.text_encoder_key_prefix[0]
-        hunyuan_detect = comfy.text_encoders.hunyuan_video.llama_detect(state_dict, "{}qwen25_7b.transformer.".format(pref))
-        return supported_models_base.ClipTarget(comfy.text_encoders.hunyuan_video.HunyuanVideo15Tokenizer, comfy.text_encoders.hunyuan_image.te(**hunyuan_detect))
-
-
-class HunyuanVideo15_SR_Distilled(HunyuanVideo):
-    unet_config = {
-        "image_model": "hunyuan_video",
-        "vision_in_dim": 1152,
-        "in_channels": 98,
-    }
-
-    sampling_settings = {
-        "shift": 2.0,
-    }
-    memory_usage_factor = 4.0 #TODO
-    supported_inference_dtypes = [torch.float16, torch.bfloat16, torch.float32]
-
-    latent_format = latent_formats.HunyuanVideo15
-
-    def get_model(self, state_dict, prefix="", device=None):
-        out = model_base.HunyuanVideo15_SR_Distilled(self, device=device)
-        return out
-
-    def clip_target(self, state_dict={}):
-        pref = self.text_encoder_key_prefix[0]
-        hunyuan_detect = comfy.text_encoders.hunyuan_video.llama_detect(state_dict, "{}qwen25_7b.transformer.".format(pref))
-        return supported_models_base.ClipTarget(comfy.text_encoders.hunyuan_video.HunyuanVideo15Tokenizer, comfy.text_encoders.hunyuan_image.te(**hunyuan_detect))
-
-models = [LotusD, Stable_Zero123, SD15_instructpix2pix, SD15, SD20, SD21UnclipL, SD21UnclipH, SDXL_instructpix2pix, SDXLRefiner, SDXL, SSD1B, KOALA_700M, KOALA_1B, Segmind_Vega, SD_X4Upscaler, Stable_Cascade_C, Stable_Cascade_B, SV3D_u, SV3D_p, SD3, StableAudio, AuraFlow, PixArtAlpha, PixArtSigma, HunyuanDiT, HunyuanDiT1, FluxInpaint, Flux, FluxSchnell, GenmoMochi, LTXV, HunyuanVideo15_SR_Distilled, HunyuanVideo15, HunyuanImage21Refiner, HunyuanImage21, HunyuanVideoSkyreelsI2V, HunyuanVideoI2V, HunyuanVideo, CosmosT2V, CosmosI2V, CosmosT2IPredict2, CosmosI2VPredict2, ZImage, Lumina2, WAN22_T2V, WAN21_T2V, WAN21_I2V, WAN21_FunControl2V, WAN21_Vace, WAN21_Camera, WAN22_Camera, WAN22_S2V, WAN21_HuMo, WAN22_Animate, Hunyuan3Dv2mini, Hunyuan3Dv2, Hunyuan3Dv2_1, HiDream, Chroma, ChromaRadiance, ACEStep, Omnigen2, QwenImage, Flux2]
-
+models = [LotusD, Stable_Zero123, SD15_instructpix2pix, SD15, SD20, SD21UnclipL, SD21UnclipH, SDXL_instructpix2pix, SDXLRefiner, SDXL, SSD1B, KOALA_700M, KOALA_1B, Segmind_Vega, SD_X4Upscaler, Stable_Cascade_C, Stable_Cascade_B, SV3D_u, SV3D_p, SD3, StableAudio, AuraFlow, PixArtAlpha, PixArtSigma, HunyuanDiT, HunyuanDiT1, FluxInpaint, Flux, FluxSchnell, GenmoMochi, LTXV, HunyuanImage21Refiner, HunyuanImage21, HunyuanVideoSkyreelsI2V, HunyuanVideoI2V, HunyuanVideo, CosmosT2V, CosmosI2V, CosmosT2IPredict2, CosmosI2VPredict2, Lumina2, WAN22_T2V, WAN21_T2V, WAN21_I2V, WAN21_FunControl2V, WAN21_Vace, WAN21_Camera, WAN22_Camera, WAN22_S2V, WAN21_HuMo, WAN22_Animate, Hunyuan3Dv2mini, Hunyuan3Dv2, Hunyuan3Dv2_1, HiDream, Chroma, ChromaRadiance, ACEStep, Omnigen2, QwenImage]

 models += [SVD_img2vid]
--- a/comfy/taesd/taehv.py
+++ b/comfy/taesd/taehv.py
@@ -1,171 +0,0 @@
-# Tiny AutoEncoder for HunyuanVideo and WanVideo https://github.com/madebyollin/taehv
-
-import torch
-import torch.nn as nn
-import torch.nn.functional as F
-from tqdm.auto import tqdm
-from collections import namedtuple, deque
-
-import comfy.ops
-operations=comfy.ops.disable_weight_init
-
-DecoderResult = namedtuple("DecoderResult", ("frame", "memory"))
-TWorkItem = namedtuple("TWorkItem", ("input_tensor", "block_index"))
-
-def conv(n_in, n_out, **kwargs):
-    return operations.Conv2d(n_in, n_out, 3, padding=1, **kwargs)
-
-class Clamp(nn.Module):
-    def forward(self, x):
-        return torch.tanh(x / 3) * 3
-
-class MemBlock(nn.Module):
-    def __init__(self, n_in, n_out, act_func):
-        super().__init__()
-        self.conv = nn.Sequential(conv(n_in * 2, n_out), act_func, conv(n_out, n_out), act_func, conv(n_out, n_out))
-        self.skip = operations.Conv2d(n_in, n_out, 1, bias=False) if n_in != n_out else nn.Identity()
-        self.act = act_func
-    def forward(self, x, past):
-        return self.act(self.conv(torch.cat([x, past], 1)) + self.skip(x))
-
-class TPool(nn.Module):
-    def __init__(self, n_f, stride):
-        super().__init__()
-        self.stride = stride
-        self.conv = operations.Conv2d(n_f*stride,n_f, 1, bias=False)
-    def forward(self, x):
-        _NT, C, H, W = x.shape
-        return self.conv(x.reshape(-1, self.stride * C, H, W))
-
-class TGrow(nn.Module):
-    def __init__(self, n_f, stride):
-        super().__init__()
-        self.stride = stride
-        self.conv = operations.Conv2d(n_f, n_f*stride, 1, bias=False)
-    def forward(self, x):
-        _NT, C, H, W = x.shape
-        x = self.conv(x)
-        return x.reshape(-1, C, H, W)
-
-def apply_model_with_memblocks(model, x, parallel, show_progress_bar):
-
-    B, T, C, H, W = x.shape
-    if parallel:
-        x = x.reshape(B*T, C, H, W)
-        # parallel over input timesteps, iterate over blocks
-        for b in tqdm(model, disable=not show_progress_bar):
-            if isinstance(b, MemBlock):
-                BT, C, H, W = x.shape
-                T = BT // B
-                _x = x.reshape(B, T, C, H, W)
-                mem = F.pad(_x, (0,0,0,0,0,0,1,0), value=0)[:,:T].reshape(x.shape)
-                x = b(x, mem)
-            else:
-                x = b(x)
-        BT, C, H, W = x.shape
-        T = BT // B
-        x = x.view(B, T, C, H, W)
-    else:
-        out = []
-        work_queue = deque([TWorkItem(xt, 0) for t, xt in enumerate(x.reshape(B, T * C, H, W).chunk(T, dim=1))])
-        progress_bar = tqdm(range(T), disable=not show_progress_bar)
-        mem = [None] * len(model)
-        while work_queue:
-            xt, i = work_queue.popleft()
-            if i == 0:
-                progress_bar.update(1)
-            if i == len(model):
-                out.append(xt)
-                del xt
-            else:
-                b = model[i]
-                if isinstance(b, MemBlock):
-                    if mem[i] is None:
-                        xt_new = b(xt, xt * 0)
-                        mem[i] = xt.detach().clone()
-                    else:
-                        xt_new = b(xt, mem[i])
-                        mem[i] = xt.detach().clone()
-                    del xt
-                    work_queue.appendleft(TWorkItem(xt_new, i+1))
-                elif isinstance(b, TPool):
-                    if mem[i] is None:
-                        mem[i] = []
-                    mem[i].append(xt.detach().clone())
-                    if len(mem[i]) == b.stride:
-                        B, C, H, W = xt.shape
-                        xt = b(torch.cat(mem[i], 1).view(B*b.stride, C, H, W))
-                        mem[i] = []
-                        work_queue.appendleft(TWorkItem(xt, i+1))
-                elif isinstance(b, TGrow):
-                    xt = b(xt)
-                    NT, C, H, W = xt.shape
-                    for xt_next in reversed(xt.view(B, b.stride*C, H, W).chunk(b.stride, 1)):
-                        work_queue.appendleft(TWorkItem(xt_next, i+1))
-                    del xt
-                else:
-                    xt = b(xt)
-                    work_queue.appendleft(TWorkItem(xt, i+1))
-        progress_bar.close()
-        x = torch.stack(out, 1)
-    return x
-
-
-class TAEHV(nn.Module):
-    def __init__(self, latent_channels, parallel=False, decoder_time_upscale=(True, True), decoder_space_upscale=(True, True, True), latent_format=None, show_progress_bar=True):
-        super().__init__()
-        self.image_channels = 3
-        self.patch_size = 1
-        self.latent_channels = latent_channels
-        self.parallel = parallel
-        self.latent_format = latent_format
-        self.show_progress_bar = show_progress_bar
-        self.process_in = latent_format().process_in if latent_format is not None else (lambda x: x)
-        self.process_out = latent_format().process_out if latent_format is not None else (lambda x: x)
-        if self.latent_channels in [48, 32]: # Wan 2.2 and HunyuanVideo1.5
-            self.patch_size = 2
-        if self.latent_channels == 32: # HunyuanVideo1.5
-            act_func = nn.LeakyReLU(0.2, inplace=True)
-        else: # HunyuanVideo, Wan 2.1
-            act_func = nn.ReLU(inplace=True)
-
-        self.encoder = nn.Sequential(
-            conv(self.image_channels*self.patch_size**2, 64), act_func,
-            TPool(64, 2), conv(64, 64, stride=2, bias=False), MemBlock(64, 64, act_func), MemBlock(64, 64, act_func), MemBlock(64, 64, act_func),
-            TPool(64, 2), conv(64, 64, stride=2, bias=False), MemBlock(64, 64, act_func), MemBlock(64, 64, act_func), MemBlock(64, 64, act_func),
-            TPool(64, 1), conv(64, 64, stride=2, bias=False), MemBlock(64, 64, act_func), MemBlock(64, 64, act_func), MemBlock(64, 64, act_func),
-            conv(64, self.latent_channels),
-        )
-        n_f = [256, 128, 64, 64]
-        self.frames_to_trim = 2**sum(decoder_time_upscale) - 1
-        self.decoder = nn.Sequential(
-            Clamp(), conv(self.latent_channels, n_f[0]), act_func,
-            MemBlock(n_f[0], n_f[0], act_func), MemBlock(n_f[0], n_f[0], act_func), MemBlock(n_f[0], n_f[0], act_func), nn.Upsample(scale_factor=2 if decoder_space_upscale[0] else 1), TGrow(n_f[0], 1), conv(n_f[0], n_f[1], bias=False),
-            MemBlock(n_f[1], n_f[1], act_func), MemBlock(n_f[1], n_f[1], act_func), MemBlock(n_f[1], n_f[1], act_func), nn.Upsample(scale_factor=2 if decoder_space_upscale[1] else 1), TGrow(n_f[1], 2 if decoder_time_upscale[0] else 1), conv(n_f[1], n_f[2], bias=False),
-            MemBlock(n_f[2], n_f[2], act_func), MemBlock(n_f[2], n_f[2], act_func), MemBlock(n_f[2], n_f[2], act_func), nn.Upsample(scale_factor=2 if decoder_space_upscale[2] else 1), TGrow(n_f[2], 2 if decoder_time_upscale[1] else 1), conv(n_f[2], n_f[3], bias=False),
-            act_func, conv(n_f[3], self.image_channels*self.patch_size**2),
-        )
-        @property
-        def show_progress_bar(self):
-            return self._show_progress_bar
-
-        @show_progress_bar.setter
-        def show_progress_bar(self, value):
-            self._show_progress_bar = value
-
-    def encode(self, x, **kwargs):
-        if self.patch_size > 1: x = F.pixel_unshuffle(x, self.patch_size)
-        x = x.movedim(2, 1)  # [B, C, T, H, W] -> [B, T, C, H, W]
-        if x.shape[1] % 4 != 0:
-            # pad at end to multiple of 4
-            n_pad = 4 - x.shape[1] % 4
-            padding = x[:, -1:].repeat_interleave(n_pad, dim=1)
-            x = torch.cat([x, padding], 1)
-        x = apply_model_with_memblocks(self.encoder, x, self.parallel, self.show_progress_bar).movedim(2, 1)
-        return self.process_out(x)
-
-    def decode(self, x, **kwargs):
-        x = self.process_in(x).movedim(2, 1)  # [B, C, T, H, W] -> [B, T, C, H, W]
-        x = apply_model_with_memblocks(self.decoder, x, self.parallel, self.show_progress_bar)
-        if self.patch_size > 1: x = F.pixel_shuffle(x, self.patch_size)
-        return x[:, self.frames_to_trim:].movedim(2, 1)
--- a/comfy/text_encoders/flux.py
+++ b/comfy/text_encoders/flux.py
@@ -1,13 +1,10 @@
 from comfy import sd1_clip
 import comfy.text_encoders.t5
 import comfy.text_encoders.sd3_clip
-import comfy.text_encoders.llama
 import comfy.model_management
-from transformers import T5TokenizerFast, LlamaTokenizerFast
+from transformers import T5TokenizerFast
 import torch
 import os
-import json
-import base64

 class T5XXLTokenizer(sd1_clip.SDTokenizer):
    def __init__(self, embedding_directory=None, tokenizer_data={}):
@@ -71,106 +68,3 @@ def flux_clip(dtype_t5=None, t5xxl_scaled_fp8=None):
                model_options["t5xxl_scaled_fp8"] = t5xxl_scaled_fp8
            super().__init__(dtype_t5=dtype_t5, device=device, dtype=dtype, model_options=model_options)
    return FluxClipModel_
-
-def load_mistral_tokenizer(data):
-    if torch.is_tensor(data):
-        data = data.numpy().tobytes()
-
-    try:
-        from transformers.integrations.mistral import MistralConverter
-    except ModuleNotFoundError:
-        from transformers.models.pixtral.convert_pixtral_weights_to_hf import MistralConverter
-
-    mistral_vocab = json.loads(data)
-
-    special_tokens = {}
-    vocab = {}
-
-    max_vocab = mistral_vocab["config"]["default_vocab_size"]
-    max_vocab -= len(mistral_vocab["special_tokens"])
-
-    for w in mistral_vocab["vocab"]:
-        r = w["rank"]
-        if r >= max_vocab:
-            continue
-
-        vocab[base64.b64decode(w["token_bytes"])] = r
-
-    for w in mistral_vocab["special_tokens"]:
-        if "token_bytes" in w:
-            special_tokens[base64.b64decode(w["token_bytes"])] = w["rank"]
-        else:
-            special_tokens[w["token_str"]] = w["rank"]
-
-    all_special = []
-    for v in special_tokens:
-        all_special.append(v)
-
-    special_tokens.update(vocab)
-    vocab = special_tokens
-    return {"tokenizer_object": MistralConverter(vocab=vocab, additional_special_tokens=all_special).converted(), "legacy": False}
-
-class MistralTokenizerClass:
-    @staticmethod
-    def from_pretrained(path, **kwargs):
-        return LlamaTokenizerFast(**kwargs)
-
-class Mistral3Tokenizer(sd1_clip.SDTokenizer):
-    def __init__(self, embedding_directory=None, tokenizer_data={}):
-        self.tekken_data = tokenizer_data.get("tekken_model", None)
-        super().__init__("", pad_with_end=False, embedding_size=5120, embedding_key='mistral3_24b', tokenizer_class=MistralTokenizerClass, has_end_token=False, pad_to_max_length=False, pad_token=11, max_length=99999999, min_length=1, pad_left=True, tokenizer_args=load_mistral_tokenizer(self.tekken_data), tokenizer_data=tokenizer_data)
-
-    def state_dict(self):
-        return {"tekken_model": self.tekken_data}
-
-class Flux2Tokenizer(sd1_clip.SD1Tokenizer):
-    def __init__(self, embedding_directory=None, tokenizer_data={}):
-        super().__init__(embedding_directory=embedding_directory, tokenizer_data=tokenizer_data, name="mistral3_24b", tokenizer=Mistral3Tokenizer)
-        self.llama_template = '[SYSTEM_PROMPT]You are an AI that reasons about image descriptions. You give structured responses focusing on object relationships, object\nattribution and actions without speculation.[/SYSTEM_PROMPT][INST]{}[/INST]'
-
-    def tokenize_with_weights(self, text, return_word_ids=False, llama_template=None, **kwargs):
-        if llama_template is None:
-            llama_text = self.llama_template.format(text)
-        else:
-            llama_text = llama_template.format(text)
-
-        tokens = super().tokenize_with_weights(llama_text, return_word_ids=return_word_ids, disable_weights=True, **kwargs)
-        return tokens
-
-class Mistral3_24BModel(sd1_clip.SDClipModel):
-    def __init__(self, device="cpu", layer=[10, 20, 30], layer_idx=None, dtype=None, attention_mask=True, model_options={}):
-        textmodel_json_config = {}
-        num_layers = model_options.get("num_layers", None)
-        if num_layers is not None:
-            textmodel_json_config["num_hidden_layers"] = num_layers
-            if num_layers < 40:
-                textmodel_json_config["final_norm"] = False
-        super().__init__(device=device, layer=layer, layer_idx=layer_idx, textmodel_json_config=textmodel_json_config, dtype=dtype, special_tokens={"start": 1, "pad": 0}, layer_norm_hidden_state=False, model_class=comfy.text_encoders.llama.Mistral3Small24B, enable_attention_masks=attention_mask, return_attention_masks=attention_mask, model_options=model_options)
-
-class Flux2TEModel(sd1_clip.SD1ClipModel):
-    def __init__(self, device="cpu", dtype=None, model_options={}, name="mistral3_24b", clip_model=Mistral3_24BModel):
-        super().__init__(device=device, dtype=dtype, name=name, clip_model=clip_model, model_options=model_options)
-
-    def encode_token_weights(self, token_weight_pairs):
-        out, pooled, extra = super().encode_token_weights(token_weight_pairs)
-
-        out = torch.stack((out[:, 0], out[:, 1], out[:, 2]), dim=1)
-        out = out.movedim(1, 2)
-        out = out.reshape(out.shape[0], out.shape[1], -1)
-        return out, pooled, extra
-
-def flux2_te(dtype_llama=None, llama_scaled_fp8=None, llama_quantization_metadata=None, pruned=False):
-    class Flux2TEModel_(Flux2TEModel):
-        def __init__(self, device="cpu", dtype=None, model_options={}):
-            if llama_scaled_fp8 is not None and "scaled_fp8" not in model_options:
-                model_options = model_options.copy()
-                model_options["scaled_fp8"] = llama_scaled_fp8
-            if dtype_llama is not None:
-                dtype = dtype_llama
-            if llama_quantization_metadata is not None:
-                model_options["quantization_metadata"] = llama_quantization_metadata
-            if pruned:
-                model_options = model_options.copy()
-                model_options["num_layers"] = 30
-            super().__init__(device=device, dtype=dtype, model_options=model_options)
-    return Flux2TEModel_
--- a/comfy/text_encoders/hunyuan_video.py
+++ b/comfy/text_encoders/hunyuan_video.py
@@ -1,7 +1,6 @@
 from comfy import sd1_clip
 import comfy.model_management
 import comfy.text_encoders.llama
-from .hunyuan_image import HunyuanImageTokenizer
 from transformers import LlamaTokenizerFast
 import torch
 import os
@@ -18,9 +17,6 @@ def llama_detect(state_dict, prefix=""):
    if scaled_fp8_key in state_dict:
        out["llama_scaled_fp8"] = state_dict[scaled_fp8_key].dtype

-    if "_quantization_metadata" in state_dict:
-        out["llama_quantization_metadata"] = state_dict["_quantization_metadata"]
-
    return out


@@ -77,14 +73,6 @@ class HunyuanVideoTokenizer:
        return {}


-class HunyuanVideo15Tokenizer(HunyuanImageTokenizer):
-    def __init__(self, embedding_directory=None, tokenizer_data={}):
-        super().__init__(embedding_directory=embedding_directory, tokenizer_data=tokenizer_data)
-        self.llama_template = "<|im_start|>system\nYou are a helpful assistant. Describe the video by detailing the following aspects:\n1. The main content and theme of the video.\n2. The color, shape, size, texture, quantity, text, and spatial relationships of the objects.\n3. Actions, events, behaviors temporal relationships, physical movement changes of the objects.\n4. background environment, light, style and atmosphere.\n5. camera angles, movements, and transitions used in the video.<|im_end|>\n<|im_start|>user\n{}<|im_end|>\n<|im_start|>assistant\n"
-
-    def tokenize_with_weights(self, text:str, return_word_ids=False, **kwargs):
-        return super().tokenize_with_weights(text, return_word_ids, prevent_empty_text=True, **kwargs)
-
 class HunyuanVideoClipModel(torch.nn.Module):
    def __init__(self, dtype_llama=None, device="cpu", dtype=None, model_options={}):
        super().__init__()
--- a/comfy/text_encoders/llama.py
+++ b/comfy/text_encoders/llama.py
@@ -32,29 +32,6 @@ class Llama2Config:
    q_norm = None
    k_norm = None
    rope_scale = None
-    final_norm: bool = True
-
-@dataclass
-class Mistral3Small24BConfig:
-    vocab_size: int = 131072
-    hidden_size: int = 5120
-    intermediate_size: int = 32768
-    num_hidden_layers: int = 40
-    num_attention_heads: int = 32
-    num_key_value_heads: int = 8
-    max_position_embeddings: int = 8192
-    rms_norm_eps: float = 1e-5
-    rope_theta: float = 1000000000.0
-    transformer_type: str = "llama"
-    head_dim = 128
-    rms_norm_add = False
-    mlp_activation = "silu"
-    qkv_bias = False
-    rope_dims = None
-    q_norm = None
-    k_norm = None
-    rope_scale = None
-    final_norm: bool = True

@dataclass
 class Qwen25_3BConfig:
@@ -76,29 +53,6 @@ class Qwen25_3BConfig:
    q_norm = None
    k_norm = None
    rope_scale = None
-    final_norm: bool = True
-
-@dataclass
-class Qwen3_4BConfig:
-    vocab_size: int = 151936
-    hidden_size: int = 2560
-    intermediate_size: int = 9728
-    num_hidden_layers: int = 36
-    num_attention_heads: int = 32
-    num_key_value_heads: int = 8
-    max_position_embeddings: int = 40960
-    rms_norm_eps: float = 1e-6
-    rope_theta: float = 1000000.0
-    transformer_type: str = "llama"
-    head_dim = 128
-    rms_norm_add = False
-    mlp_activation = "silu"
-    qkv_bias = False
-    rope_dims = None
-    q_norm = "gemma3"
-    k_norm = "gemma3"
-    rope_scale = None
-    final_norm: bool = True

@dataclass
 class Qwen25_7BVLI_Config:
@@ -120,7 +74,6 @@ class Qwen25_7BVLI_Config:
    q_norm = None
    k_norm = None
    rope_scale = None
-    final_norm: bool = True

@dataclass
 class Gemma2_2B_Config:
@@ -143,7 +96,6 @@ class Gemma2_2B_Config:
    k_norm = None
    sliding_attention = None
    rope_scale = None
-    final_norm: bool = True

@dataclass
 class Gemma3_4B_Config:
@@ -166,7 +118,6 @@ class Gemma3_4B_Config:
    k_norm = "gemma3"
    sliding_attention = [False, False, False, False, False, 1024]
    rope_scale = [1.0, 8.0]
-    final_norm: bool = True

 class RMSNorm(nn.Module):
    def __init__(self, dim: int, eps: float = 1e-5, add=False, device=None, dtype=None):
@@ -415,12 +366,7 @@ class Llama2_(nn.Module):
            transformer(config, index=i, device=device, dtype=dtype, ops=ops)
            for i in range(config.num_hidden_layers)
        ])
-
-        if config.final_norm:
-            self.norm = RMSNorm(config.hidden_size, eps=config.rms_norm_eps, add=config.rms_norm_add, device=device, dtype=dtype)
-        else:
-            self.norm = None
-
+        self.norm = RMSNorm(config.hidden_size, eps=config.rms_norm_eps, add=config.rms_norm_add, device=device, dtype=dtype)
        # self.lm_head = ops.Linear(config.hidden_size, config.vocab_size, bias=False, device=device, dtype=dtype)

    def forward(self, x, attention_mask=None, embeds=None, num_tokens=None, intermediate_output=None, final_layer_norm_intermediate=True, dtype=None, position_ids=None, embeds_info=[]):
@@ -456,12 +402,8 @@ class Llama2_(nn.Module):

        intermediate = None
        all_intermediate = None
-        only_layers = None
        if intermediate_output is not None:
-            if isinstance(intermediate_output, list):
-                all_intermediate = []
-                only_layers = set(intermediate_output)
-            elif intermediate_output == "all":
+            if intermediate_output == "all":
                all_intermediate = []
                intermediate_output = None
            elif intermediate_output < 0:
@@ -469,8 +411,7 @@ class Llama2_(nn.Module):

        for i, layer in enumerate(self.layers):
            if all_intermediate is not None:
-                if only_layers is None or (i in only_layers):
-                    all_intermediate.append(x.unsqueeze(1).clone())
+                all_intermediate.append(x.unsqueeze(1).clone())
            x = layer(
                x=x,
                attention_mask=mask,
@@ -480,17 +421,14 @@ class Llama2_(nn.Module):
            if i == intermediate_output:
                intermediate = x.clone()

-        if self.norm is not None:
-            x = self.norm(x)
-
+        x = self.norm(x)
        if all_intermediate is not None:
-            if only_layers is None or ((i + 1) in only_layers):
-                all_intermediate.append(x.unsqueeze(1).clone())
+            all_intermediate.append(x.unsqueeze(1).clone())

        if all_intermediate is not None:
            intermediate = torch.cat(all_intermediate, dim=1)

-        if intermediate is not None and final_layer_norm_intermediate and self.norm is not None:
+        if intermediate is not None and final_layer_norm_intermediate:
            intermediate = self.norm(intermediate)

        return x, intermediate
@@ -515,15 +453,6 @@ class Llama2(BaseLlama, torch.nn.Module):
        self.model = Llama2_(config, device=device, dtype=dtype, ops=operations)
        self.dtype = dtype

-class Mistral3Small24B(BaseLlama, torch.nn.Module):
-    def __init__(self, config_dict, dtype, device, operations):
-        super().__init__()
-        config = Mistral3Small24BConfig(**config_dict)
-        self.num_layers = config.num_hidden_layers
-
-        self.model = Llama2_(config, device=device, dtype=dtype, ops=operations)
-        self.dtype = dtype
-
 class Qwen25_3B(BaseLlama, torch.nn.Module):
    def __init__(self, config_dict, dtype, device, operations):
        super().__init__()
@@ -533,15 +462,6 @@ class Qwen25_3B(BaseLlama, torch.nn.Module):
        self.model = Llama2_(config, device=device, dtype=dtype, ops=operations)
        self.dtype = dtype

-class Qwen3_4B(BaseLlama, torch.nn.Module):
-    def __init__(self, config_dict, dtype, device, operations):
-        super().__init__()
-        config = Qwen3_4BConfig(**config_dict)
-        self.num_layers = config.num_hidden_layers
-
-        self.model = Llama2_(config, device=device, dtype=dtype, ops=operations)
-        self.dtype = dtype
-
 class Qwen25_7BVLI(BaseLlama, torch.nn.Module):
    def __init__(self, config_dict, dtype, device, operations):
        super().__init__()
--- a/comfy/text_encoders/qwen25_tokenizer/tokenizer_config.json
+++ b/comfy/text_encoders/qwen25_tokenizer/tokenizer_config.json
@@ -179,36 +179,36 @@
      "special": false
    },
    "151665": {
-      "content": "<tool_response>",
+      "content": "<|img|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
-      "special": false
+      "special": true
    },
    "151666": {
-      "content": "</tool_response>",
+      "content": "<|endofimg|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
-      "special": false
+      "special": true
    },
    "151667": {
-      "content": "<think>",
+      "content": "<|meta|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
-      "special": false
+      "special": true
    },
    "151668": {
-      "content": "</think>",
+      "content": "<|endofmeta|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
-      "special": false
+      "special": true
    }
  },
  "additional_special_tokens": [
--- a/comfy/text_encoders/qwen_image.py
+++ b/comfy/text_encoders/qwen_image.py
@@ -17,14 +17,12 @@ class QwenImageTokenizer(sd1_clip.SD1Tokenizer):
        self.llama_template = "<|im_start|>system\nDescribe the image by detailing the color, shape, size, texture, quantity, text, spatial relationships of the objects and background:<|im_end|>\n<|im_start|>user\n{}<|im_end|>\n<|im_start|>assistant\n"
        self.llama_template_images = "<|im_start|>system\nDescribe the key features of the input image (color, shape, size, texture, objects, background), then explain how the user's text instruction should alter or modify the image. Generate a new image that meets the user's requirements while maintaining consistency with the original input where appropriate.<|im_end|>\n<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|>{}<|im_end|>\n<|im_start|>assistant\n"

-    def tokenize_with_weights(self, text, return_word_ids=False, llama_template=None, images=[], prevent_empty_text=False, **kwargs):
+    def tokenize_with_weights(self, text, return_word_ids=False, llama_template=None, images=[], **kwargs):
        skip_template = False
        if text.startswith('<|im_start|>'):
            skip_template = True
        if text.startswith('<|start_header_id|>'):
            skip_template = True
-        if prevent_empty_text and text == '':
-            text = ' '

        if skip_template:
            llama_text = text
--- a/comfy/text_encoders/z_image.py
+++ b/comfy/text_encoders/z_image.py
@@ -1,48 +0,0 @@
-from transformers import Qwen2Tokenizer
-import comfy.text_encoders.llama
-from comfy import sd1_clip
-import os
-
-class Qwen3Tokenizer(sd1_clip.SDTokenizer):
-    def __init__(self, embedding_directory=None, tokenizer_data={}):
-        tokenizer_path = os.path.join(os.path.dirname(os.path.realpath(__file__)), "qwen25_tokenizer")
-        super().__init__(tokenizer_path, pad_with_end=False, embedding_size=2560, embedding_key='qwen3_4b', tokenizer_class=Qwen2Tokenizer, has_start_token=False, has_end_token=False, pad_to_max_length=False, max_length=99999999, min_length=1, pad_token=151643, tokenizer_data=tokenizer_data)
-
-
-class ZImageTokenizer(sd1_clip.SD1Tokenizer):
-    def __init__(self, embedding_directory=None, tokenizer_data={}):
-        super().__init__(embedding_directory=embedding_directory, tokenizer_data=tokenizer_data, name="qwen3_4b", tokenizer=Qwen3Tokenizer)
-        self.llama_template = "<|im_start|>user\n{}<|im_end|>\n<|im_start|>assistant\n"
-
-    def tokenize_with_weights(self, text, return_word_ids=False, llama_template=None, **kwargs):
-        if llama_template is None:
-            llama_text = self.llama_template.format(text)
-        else:
-            llama_text = llama_template.format(text)
-
-        tokens = super().tokenize_with_weights(llama_text, return_word_ids=return_word_ids, disable_weights=True, **kwargs)
-        return tokens
-
-
-class Qwen3_4BModel(sd1_clip.SDClipModel):
-    def __init__(self, device="cpu", layer="hidden", layer_idx=-2, dtype=None, attention_mask=True, model_options={}):
-        super().__init__(device=device, layer=layer, layer_idx=layer_idx, textmodel_json_config={}, dtype=dtype, special_tokens={"pad": 151643}, layer_norm_hidden_state=False, model_class=comfy.text_encoders.llama.Qwen3_4B, enable_attention_masks=attention_mask, return_attention_masks=attention_mask, model_options=model_options)
-
-
-class ZImageTEModel(sd1_clip.SD1ClipModel):
-    def __init__(self, device="cpu", dtype=None, model_options={}):
-        super().__init__(device=device, dtype=dtype, name="qwen3_4b", clip_model=Qwen3_4BModel, model_options=model_options)
-
-
-def te(dtype_llama=None, llama_scaled_fp8=None, llama_quantization_metadata=None):
-    class ZImageTEModel_(ZImageTEModel):
-        def __init__(self, device="cpu", dtype=None, model_options={}):
-            if llama_scaled_fp8 is not None and "scaled_fp8" not in model_options:
-                model_options = model_options.copy()
-                model_options["scaled_fp8"] = llama_scaled_fp8
-            if dtype_llama is not None:
-                dtype = dtype_llama
-            if llama_quantization_metadata is not None:
-                model_options["quantization_metadata"] = llama_quantization_metadata
-            super().__init__(device=device, dtype=dtype, model_options=model_options)
-    return ZImageTEModel_
--- a/comfy/utils.py
+++ b/comfy/utils.py
@@ -675,72 +675,6 @@ def flux_to_diffusers(mmdit_config, output_prefix=""):

    return key_map

-def z_image_to_diffusers(mmdit_config, output_prefix=""):
-    n_layers = mmdit_config.get("n_layers", 0)
-    hidden_size = mmdit_config.get("dim", 0)
-    n_context_refiner = mmdit_config.get("n_refiner_layers", 2)
-    n_noise_refiner = mmdit_config.get("n_refiner_layers", 2)
-    key_map = {}
-
-    def add_block_keys(prefix_from, prefix_to, has_adaln=True):
-        for end in ("weight", "bias"):
-            k = "{}.attention.".format(prefix_from)
-            qkv = "{}.attention.qkv.{}".format(prefix_to, end)
-            key_map["{}to_q.{}".format(k, end)] = (qkv, (0, 0, hidden_size))
-            key_map["{}to_k.{}".format(k, end)] = (qkv, (0, hidden_size, hidden_size))
-            key_map["{}to_v.{}".format(k, end)] = (qkv, (0, hidden_size * 2, hidden_size))
-
-        block_map = {
-            "attention.norm_q.weight": "attention.q_norm.weight",
-            "attention.norm_k.weight": "attention.k_norm.weight",
-            "attention.to_out.0.weight": "attention.out.weight",
-            "attention.to_out.0.bias": "attention.out.bias",
-            "attention_norm1.weight": "attention_norm1.weight",
-            "attention_norm2.weight": "attention_norm2.weight",
-            "feed_forward.w1.weight": "feed_forward.w1.weight",
-            "feed_forward.w2.weight": "feed_forward.w2.weight",
-            "feed_forward.w3.weight": "feed_forward.w3.weight",
-            "ffn_norm1.weight": "ffn_norm1.weight",
-            "ffn_norm2.weight": "ffn_norm2.weight",
-        }
-        if has_adaln:
-            block_map["adaLN_modulation.0.weight"] = "adaLN_modulation.0.weight"
-            block_map["adaLN_modulation.0.bias"] = "adaLN_modulation.0.bias"
-        for k, v in block_map.items():
-            key_map["{}.{}".format(prefix_from, k)] = "{}.{}".format(prefix_to, v)
-
-    for i in range(n_layers):
-        add_block_keys("layers.{}".format(i), "{}layers.{}".format(output_prefix, i))
-
-    for i in range(n_context_refiner):
-        add_block_keys("context_refiner.{}".format(i), "{}context_refiner.{}".format(output_prefix, i))
-
-    for i in range(n_noise_refiner):
-        add_block_keys("noise_refiner.{}".format(i), "{}noise_refiner.{}".format(output_prefix, i))
-
-    MAP_BASIC = [
-        ("final_layer.linear.weight", "all_final_layer.2-1.linear.weight"),
-        ("final_layer.linear.bias", "all_final_layer.2-1.linear.bias"),
-        ("final_layer.adaLN_modulation.1.weight", "all_final_layer.2-1.adaLN_modulation.1.weight"),
-        ("final_layer.adaLN_modulation.1.bias", "all_final_layer.2-1.adaLN_modulation.1.bias"),
-        ("x_embedder.weight", "all_x_embedder.2-1.weight"),
-        ("x_embedder.bias", "all_x_embedder.2-1.bias"),
-        ("x_pad_token", "x_pad_token"),
-        ("cap_embedder.0.weight", "cap_embedder.0.weight"),
-        ("cap_embedder.1.weight", "cap_embedder.1.weight"),
-        ("cap_embedder.1.bias", "cap_embedder.1.bias"),
-        ("cap_pad_token", "cap_pad_token"),
-        ("t_embedder.mlp.0.weight", "t_embedder.mlp.0.weight"),
-        ("t_embedder.mlp.0.bias", "t_embedder.mlp.0.bias"),
-        ("t_embedder.mlp.2.weight", "t_embedder.mlp.2.weight"),
-        ("t_embedder.mlp.2.bias", "t_embedder.mlp.2.bias"),
-    ]
-
-    for c, diffusers in MAP_BASIC:
-        key_map[diffusers] = "{}{}".format(output_prefix, c)
-
-    return key_map
-
 def repeat_to_batch_size(tensor, batch_size, dim=0):
    if tensor.shape[dim] > batch_size:
        return tensor.narrow(dim, 0, batch_size)
--- a/comfy/weight_adapter/lora.py
+++ b/comfy/weight_adapter/lora.py
@@ -194,7 +194,6 @@ class LoRAAdapter(WeightAdapterBase):
            lora_diff = torch.mm(
                mat1.flatten(start_dim=1), mat2.flatten(start_dim=1)
            ).reshape(weight.shape)
-            del mat1, mat2
            if dora_scale is not None:
                weight = weight_decompose(
                    dora_scale,
--- a/comfy_api/internal/async_to_sync.py
+++ b/comfy_api/internal/async_to_sync.py
@@ -8,7 +8,7 @@ import os
 import textwrap
 import threading
 from enum import Enum
-from typing import Optional, Type, get_origin, get_args, get_type_hints
+from typing import Optional, Type, get_origin, get_args


 class TypeTracker:
@@ -220,18 +220,11 @@ class AsyncToSyncConverter:
            self._async_instance = async_class(*args, **kwargs)

            # Handle annotated class attributes (like execution: Execution)
-            # Get all annotations from the class hierarchy and resolve string annotations
-            try:
-                # get_type_hints resolves string annotations to actual type objects
-                # This handles classes using 'from __future__ import annotations'
-                all_annotations = get_type_hints(async_class)
-            except Exception:
-                # Fallback to raw annotations if get_type_hints fails
-                # (e.g., for undefined forward references)
-                all_annotations = {}
-                for base_class in reversed(inspect.getmro(async_class)):
-                    if hasattr(base_class, "__annotations__"):
-                        all_annotations.update(base_class.__annotations__)
+            # Get all annotations from the class hierarchy
+            all_annotations = {}
+            for base_class in reversed(inspect.getmro(async_class)):
+                if hasattr(base_class, "__annotations__"):
+                    all_annotations.update(base_class.__annotations__)

            # For each annotated attribute, check if it needs to be created or wrapped
            for attr_name, attr_type in all_annotations.items():
@@ -632,19 +625,15 @@ class AsyncToSyncConverter:
        """Extract class attributes that are classes themselves."""
        class_attributes = []

-        # Get resolved type hints to handle string annotations
-        try:
-            type_hints = get_type_hints(async_class)
-        except Exception:
-            type_hints = {}
-
        # Look for class attributes that are classes
        for name, attr in sorted(inspect.getmembers(async_class)):
            if isinstance(attr, type) and not name.startswith("_"):
                class_attributes.append((name, attr))
-            elif name in type_hints:
-                # Use resolved type hint instead of raw annotation
-                annotation = type_hints[name]
+            elif (
+                hasattr(async_class, "__annotations__")
+                and name in async_class.__annotations__
+            ):
+                annotation = async_class.__annotations__[name]
                if isinstance(annotation, type):
                    class_attributes.append((name, annotation))

@@ -919,15 +908,11 @@ class AsyncToSyncConverter:
            attribute_mappings = {}

            # First check annotations for typed attributes (including from parent classes)
-            # Resolve string annotations to actual types
-            try:
-                all_annotations = get_type_hints(async_class)
-            except Exception:
-                # Fallback to raw annotations
-                all_annotations = {}
-                for base_class in reversed(inspect.getmro(async_class)):
-                    if hasattr(base_class, "__annotations__"):
-                        all_annotations.update(base_class.__annotations__)
+            # Collect all annotations from the class hierarchy
+            all_annotations = {}
+            for base_class in reversed(inspect.getmro(async_class)):
+                if hasattr(base_class, "__annotations__"):
+                    all_annotations.update(base_class.__annotations__)

            for attr_name, attr_type in sorted(all_annotations.items()):
                for class_name, class_type in class_attributes:
--- a/comfy_api/latest/init.py
+++ b/comfy_api/latest/init.py
@@ -7,7 +7,7 @@ from comfy_api.internal.singleton import ProxiedSingleton
 from comfy_api.internal.async_to_sync import create_sync_class
 from comfy_api.latest._input import ImageInput, AudioInput, MaskInput, LatentInput, VideoInput
 from comfy_api.latest._input_impl import VideoFromFile, VideoFromComponents
-from comfy_api.latest._util import VideoCodec, VideoContainer, VideoComponents, MESH, VOXEL
+from comfy_api.latest._util import VideoCodec, VideoContainer, VideoComponents
 from . import _io as io
 from . import _ui as ui
 # from comfy_api.latest._resources import _RESOURCES as resources  #noqa: F401
@@ -104,8 +104,6 @@ class Types:
    VideoCodec = VideoCodec
    VideoContainer = VideoContainer
    VideoComponents = VideoComponents
-    MESH = MESH
-    VOXEL = VOXEL

 ComfyAPI = ComfyAPI_latest

--- a/comfy_api/latest/_input/video_types.py
+++ b/comfy_api/latest/_input/video_types.py
@@ -1,6 +1,5 @@
 from __future__ import annotations
 from abc import ABC, abstractmethod
-from fractions import Fraction
 from typing import Optional, Union, IO
 import io
 import av
@@ -73,33 +72,6 @@ class VideoInput(ABC):
        frame_count = components.images.shape[0]
        return float(frame_count / components.frame_rate)

-    def get_frame_count(self) -> int:
-        """
-        Returns the number of frames in the video.
-
-        Default implementation uses :meth:`get_components`, which may require
-        loading all frames into memory. File-based implementations should
-        override this method and use container/stream metadata instead.
-
-        Returns:
-            Total number of frames as an integer.
-        """
-        return int(self.get_components().images.shape[0])
-
-    def get_frame_rate(self) -> Fraction:
-        """
-        Returns the frame rate of the video.
-
-        Default implementation materializes the video into memory via
-        `get_components()`. Subclasses that can inspect the underlying
-        container (e.g. `VideoFromFile`) should override this with a more
-        efficient implementation.
-
-        Returns:
-            Frame rate as a Fraction.
-        """
-        return self.get_components().frame_rate
-
    def get_container_format(self) -> str:
        """
        Returns the container format of the video (e.g., 'mp4', 'mov', 'avi').
--- a/comfy_api/latest/_input_impl/video_types.py
+++ b/comfy_api/latest/_input_impl/video_types.py
@@ -121,71 +121,6 @@ class VideoFromFile(VideoInput):

        raise ValueError(f"Could not determine duration for file '{self.__file}'")

-    def get_frame_count(self) -> int:
-        """
-        Returns the number of frames in the video without materializing them as
-        torch tensors.
-        """
-        if isinstance(self.__file, io.BytesIO):
-            self.__file.seek(0)
-
-        with av.open(self.__file, mode="r") as container:
-            video_stream = self._get_first_video_stream(container)
-            # 1. Prefer the frames field if available
-            if video_stream.frames and video_stream.frames > 0:
-                return int(video_stream.frames)
-
-            # 2. Try to estimate from duration and average_rate using only metadata
-            if container.duration is not None and video_stream.average_rate:
-                duration_seconds = float(container.duration / av.time_base)
-                estimated_frames = int(round(duration_seconds * float(video_stream.average_rate)))
-                if estimated_frames > 0:
-                    return estimated_frames
-
-            if (
-                getattr(video_stream, "duration", None) is not None
-                and getattr(video_stream, "time_base", None) is not None
-                and video_stream.average_rate
-            ):
-                duration_seconds = float(video_stream.duration * video_stream.time_base)
-                estimated_frames = int(round(duration_seconds * float(video_stream.average_rate)))
-                if estimated_frames > 0:
-                    return estimated_frames
-
-            # 3. Last resort: decode frames and count them (streaming)
-            frame_count = 0
-            container.seek(0)
-            for packet in container.demux(video_stream):
-                for _ in packet.decode():
-                    frame_count += 1
-
-            if frame_count == 0:
-                raise ValueError(f"Could not determine frame count for file '{self.__file}'")
-            return frame_count
-
-    def get_frame_rate(self) -> Fraction:
-        """
-        Returns the average frame rate of the video using container metadata
-        without decoding all frames.
-        """
-        if isinstance(self.__file, io.BytesIO):
-            self.__file.seek(0)
-
-        with av.open(self.__file, mode="r") as container:
-            video_stream = self._get_first_video_stream(container)
-            # Preferred: use PyAV's average_rate (usually already a Fraction-like)
-            if video_stream.average_rate:
-                return Fraction(video_stream.average_rate)
-
-            # Fallback: estimate from frames + duration if available
-            if video_stream.frames and container.duration:
-                duration_seconds = float(container.duration / av.time_base)
-                if duration_seconds > 0:
-                    return Fraction(video_stream.frames / duration_seconds).limit_denominator()
-
-            # Last resort: match get_components_internal default
-            return Fraction(1)
-
    def get_container_format(self) -> str:
        """
        Returns the container format of the video (e.g., 'mp4', 'mov', 'avi').
@@ -303,13 +238,6 @@ class VideoFromFile(VideoInput):
                        packet.stream = stream_map[packet.stream]
                        output_container.mux(packet)

-    def _get_first_video_stream(self, container: InputContainer):
-        video_stream = next((s for s in container.streams if s.type == "video"), None)
-        if video_stream is None:
-            raise ValueError(f"No video stream found in file '{self.__file}'")
-        return video_stream
-
-
 class VideoFromComponents(VideoInput):
    """
    Class representing video input from tensors.
@@ -336,10 +264,7 @@ class VideoFromComponents(VideoInput):
            raise ValueError("Only MP4 format is supported for now")
        if codec != VideoCodec.AUTO and codec != VideoCodec.H264:
            raise ValueError("Only H264 codec is supported for now")
-        extra_kwargs = {}
-        if isinstance(format, VideoContainer) and format != VideoContainer.AUTO:
-            extra_kwargs["format"] = format.value
-        with av.open(path, mode='w', options={'movflags': 'use_metadata_tags'}, **extra_kwargs) as output:
+        with av.open(path, mode='w', options={'movflags': 'use_metadata_tags'}) as output:
            # Add metadata before writing any streams
            if metadata is not None:
                for key, value in metadata.items():
--- a/comfy_api/latest/_io.py
+++ b/comfy_api/latest/_io.py
@@ -27,7 +27,6 @@ from comfy_api.internal import (_ComfyNodeInternal, _NodeOutputInternal, classpr
    prune_dict, shallow_clone_class)
 from comfy_api.latest._resources import Resources, ResourcesLocal
 from comfy_execution.graph_utils import ExecutionBlocker
-from ._util import MESH, VOXEL

 # from comfy_extras.nodes_images import SVG as SVG_ # NOTE: needs to be moved before can be imported due to circular reference

@@ -629,10 +628,6 @@ class UpscaleModel(ComfyTypeIO):
    if TYPE_CHECKING:
        Type = ImageModelDescriptor

-@comfytype(io_type="LATENT_UPSCALE_MODEL")
-class LatentUpscaleModel(ComfyTypeIO):
-    Type = Any
-
@comfytype(io_type="AUDIO")
 class Audio(ComfyTypeIO):
    class AudioDict(TypedDict):
@@ -661,11 +656,11 @@ class LossMap(ComfyTypeIO):

@comfytype(io_type="VOXEL")
 class Voxel(ComfyTypeIO):
-    Type = VOXEL
+    Type = Any # TODO: VOXEL class is defined in comfy_extras/nodes_hunyuan3d.py; should be moved to somewhere else before referenced directly in v3

@comfytype(io_type="MESH")
 class Mesh(ComfyTypeIO):
-    Type = MESH
+    Type = Any # TODO: MESH class is defined in comfy_extras/nodes_hunyuan3d.py; should be moved to somewhere else before referenced directly in v3

@comfytype(io_type="HOOKS")
 class Hooks(ComfyTypeIO):
--- a/comfy_api/latest/_util/init.py
+++ b/comfy_api/latest/_util/init.py
@@ -1,11 +1,8 @@
 from .video_types import VideoContainer, VideoCodec, VideoComponents
-from .geometry_types import VOXEL, MESH

 __all__ = [
    # Utility Types
    "VideoContainer",
    "VideoCodec",
    "VideoComponents",
-    "VOXEL",
-    "MESH",
 ]
--- a/comfy_api/latest/_util/geometry_types.py
+++ b/comfy_api/latest/_util/geometry_types.py
@@ -1,12 +0,0 @@
-import torch
-
-
-class VOXEL:
-    def __init__(self, data: torch.Tensor):
-        self.data = data
-
-
-class MESH:
-    def __init__(self, vertices: torch.Tensor, faces: torch.Tensor):
-        self.vertices = vertices
-        self.faces = faces
--- a/comfy_api_nodes/apinode_utils.py
+++ b/comfy_api_nodes/apinode_utils.py
@@ -0,0 +1,261 @@
+from __future__ import annotations
+import aiohttp
+import mimetypes
+from typing import Optional, Union
+from comfy.utils import common_upscale
+from comfy_api_nodes.apis.client import (
+    ApiClient,
+    ApiEndpoint,
+    HttpMethod,
+    SynchronousOperation,
+    UploadRequest,
+    UploadResponse,
+)
+from server import PromptServer
+from comfy.cli_args import args
+
+import numpy as np
+from PIL import Image
+import torch
+import math
+import base64
+from .util import tensor_to_bytesio, bytesio_to_image_tensor
+from io import BytesIO
+
+
+async def validate_and_cast_response(
+    response, timeout: int = None, node_id: Union[str, None] = None
+) -> torch.Tensor:
+    """Validates and casts a response to a torch.Tensor.
+
+    Args:
+        response: The response to validate and cast.
+        timeout: Request timeout in seconds. Defaults to None (no timeout).
+
+    Returns:
+        A torch.Tensor representing the image (1, H, W, C).
+
+    Raises:
+        ValueError: If the response is not valid.
+    """
+    # validate raw JSON response
+    data = response.data
+    if not data or len(data) == 0:
+        raise ValueError("No images returned from API endpoint")
+
+    # Initialize list to store image tensors
+    image_tensors: list[torch.Tensor] = []
+
+    # Process each image in the data array
+    async with aiohttp.ClientSession(timeout=aiohttp.ClientTimeout(total=timeout)) as session:
+        for img_data in data:
+            img_bytes: bytes
+            if img_data.b64_json:
+                img_bytes = base64.b64decode(img_data.b64_json)
+            elif img_data.url:
+                if node_id:
+                    PromptServer.instance.send_progress_text(f"Result URL: {img_data.url}", node_id)
+                async with session.get(img_data.url) as resp:
+                    if resp.status != 200:
+                        raise ValueError("Failed to download generated image")
+                    img_bytes = await resp.read()
+            else:
+                raise ValueError("Invalid image payload – neither URL nor base64 data present.")
+
+            pil_img = Image.open(BytesIO(img_bytes)).convert("RGBA")
+            arr = np.asarray(pil_img).astype(np.float32) / 255.0
+            image_tensors.append(torch.from_numpy(arr))
+
+    return torch.stack(image_tensors, dim=0)
+
+
+def validate_aspect_ratio(
+    aspect_ratio: str,
+    minimum_ratio: float,
+    maximum_ratio: float,
+    minimum_ratio_str: str,
+    maximum_ratio_str: str,
+) -> float:
+    """Validates and casts an aspect ratio string to a float.
+
+    Args:
+        aspect_ratio: The aspect ratio string to validate.
+        minimum_ratio: The minimum aspect ratio.
+        maximum_ratio: The maximum aspect ratio.
+        minimum_ratio_str: The minimum aspect ratio string.
+        maximum_ratio_str: The maximum aspect ratio string.
+
+    Returns:
+        The validated and cast aspect ratio.
+
+    Raises:
+        Exception: If the aspect ratio is not valid.
+    """
+    # get ratio values
+    numbers = aspect_ratio.split(":")
+    if len(numbers) != 2:
+        raise TypeError(
+            f"Aspect ratio must be in the format X:Y, such as 16:9, but was {aspect_ratio}."
+        )
+    try:
+        numerator = int(numbers[0])
+        denominator = int(numbers[1])
+    except ValueError as exc:
+        raise TypeError(
+            f"Aspect ratio must contain numbers separated by ':', such as 16:9, but was {aspect_ratio}."
+        ) from exc
+    calculated_ratio = numerator / denominator
+    # if not close to minimum and maximum, check bounds
+    if not math.isclose(calculated_ratio, minimum_ratio) or not math.isclose(
+        calculated_ratio, maximum_ratio
+    ):
+        if calculated_ratio < minimum_ratio:
+            raise TypeError(
+                f"Aspect ratio cannot reduce to any less than {minimum_ratio_str} ({minimum_ratio}), but was {aspect_ratio} ({calculated_ratio})."
+            )
+        if calculated_ratio > maximum_ratio:
+            raise TypeError(
+                f"Aspect ratio cannot reduce to any greater than {maximum_ratio_str} ({maximum_ratio}), but was {aspect_ratio} ({calculated_ratio})."
+            )
+    return aspect_ratio
+
+
+async def download_url_to_bytesio(
+    url: str, timeout: int = None, auth_kwargs: Optional[dict[str, str]] = None
+) -> BytesIO:
+    """Downloads content from a URL using requests and returns it as BytesIO.
+
+    Args:
+        url: The URL to download.
+        timeout: Request timeout in seconds. Defaults to None (no timeout).
+
+    Returns:
+        BytesIO object containing the downloaded content.
+    """
+    headers = {}
+    if url.startswith("/proxy/"):
+        url = str(args.comfy_api_base).rstrip("/") + url
+        auth_token = auth_kwargs.get("auth_token")
+        comfy_api_key = auth_kwargs.get("comfy_api_key")
+        if auth_token:
+            headers["Authorization"] = f"Bearer {auth_token}"
+        elif comfy_api_key:
+            headers["X-API-KEY"] = comfy_api_key
+    timeout_cfg = aiohttp.ClientTimeout(total=timeout) if timeout else None
+    async with aiohttp.ClientSession(timeout=timeout_cfg) as session:
+        async with session.get(url, headers=headers) as resp:
+            resp.raise_for_status()  # Raises HTTPError for bad responses (4XX or 5XX)
+            return BytesIO(await resp.read())
+
+
+def process_image_response(response_content: bytes | str) -> torch.Tensor:
+    """Uses content from a Response object and converts it to a torch.Tensor"""
+    return bytesio_to_image_tensor(BytesIO(response_content))
+
+
+def text_filepath_to_base64_string(filepath: str) -> str:
+    """Converts a text file to a base64 string."""
+    with open(filepath, "rb") as f:
+        file_content = f.read()
+    return base64.b64encode(file_content).decode("utf-8")
+
+
+def text_filepath_to_data_uri(filepath: str) -> str:
+    """Converts a text file to a data URI."""
+    base64_string = text_filepath_to_base64_string(filepath)
+    mime_type, _ = mimetypes.guess_type(filepath)
+    if mime_type is None:
+        mime_type = "application/octet-stream"
+    return f"data:{mime_type};base64,{base64_string}"
+
+
+async def upload_file_to_comfyapi(
+    file_bytes_io: BytesIO,
+    filename: str,
+    upload_mime_type: Optional[str],
+    auth_kwargs: Optional[dict[str, str]] = None,
+) -> str:
+    """
+    Uploads a single file to ComfyUI API and returns its download URL.
+
+    Args:
+        file_bytes_io: BytesIO object containing the file data.
+        filename: The filename of the file.
+        upload_mime_type: MIME type of the file.
+        auth_kwargs: Optional authentication token(s).
+
+    Returns:
+        The download URL for the uploaded file.
+    """
+    if upload_mime_type is None:
+        request_object = UploadRequest(file_name=filename)
+    else:
+        request_object = UploadRequest(file_name=filename, content_type=upload_mime_type)
+    operation = SynchronousOperation(
+        endpoint=ApiEndpoint(
+            path="/customers/storage",
+            method=HttpMethod.POST,
+            request_model=UploadRequest,
+            response_model=UploadResponse,
+        ),
+        request=request_object,
+        auth_kwargs=auth_kwargs,
+    )
+
+    response: UploadResponse = await operation.execute()
+    await ApiClient.upload_file(response.upload_url, file_bytes_io, content_type=upload_mime_type)
+    return response.download_url
+
+
+async def upload_images_to_comfyapi(
+    image: torch.Tensor,
+    max_images=8,
+    auth_kwargs: Optional[dict[str, str]] = None,
+    mime_type: Optional[str] = None,
+) -> list[str]:
+    """
+    Uploads images to ComfyUI API and returns download URLs.
+    To upload multiple images, stack them in the batch dimension first.
+
+    Args:
+        image: Input torch.Tensor image.
+        max_images: Maximum number of images to upload.
+        auth_kwargs: Optional authentication token(s).
+        mime_type: Optional MIME type for the image.
+    """
+    # if batch, try to upload each file if max_images is greater than 0
+    download_urls: list[str] = []
+    is_batch = len(image.shape) > 3
+    batch_len = image.shape[0] if is_batch else 1
+
+    for idx in range(min(batch_len, max_images)):
+        tensor = image[idx] if is_batch else image
+        img_io = tensor_to_bytesio(tensor, mime_type=mime_type)
+        url = await upload_file_to_comfyapi(img_io, img_io.name, mime_type, auth_kwargs)
+        download_urls.append(url)
+    return download_urls
+
+
+def resize_mask_to_image(
+    mask: torch.Tensor,
+    image: torch.Tensor,
+    upscale_method="nearest-exact",
+    crop="disabled",
+    allow_gradient=True,
+    add_channel_dim=False,
+):
+    """
+    Resize mask to be the same dimensions as an image, while maintaining proper format for API calls.
+    """
+    _, H, W, _ = image.shape
+    mask = mask.unsqueeze(-1)
+    mask = mask.movedim(-1, 1)
+    mask = common_upscale(
+        mask, width=W, height=H, upscale_method=upscale_method, crop=crop
+    )
+    mask = mask.movedim(1, -1)
+    if not add_channel_dim:
+        mask = mask.squeeze(-1)
+    if not allow_gradient:
+        mask = (mask > 0.5).float()
+    return mask
--- a/comfy_api_nodes/apis/PixverseController.py
+++ b/comfy_api_nodes/apis/PixverseController.py
@@ -0,0 +1,17 @@
+# generated by datamodel-codegen:
+#   filename:  filtered-openapi.yaml
+#   timestamp: 2025-04-29T23:44:54+00:00
+
+from __future__ import annotations
+
+from typing import Optional
+
+from pydantic import BaseModel
+
+from . import PixverseDto
+
+
+class ResponseData(BaseModel):
+    ErrCode: Optional[int] = None
+    ErrMsg: Optional[str] = None
+    Resp: Optional[PixverseDto.V2OpenAPII2VResp] = None
--- a/comfy_api_nodes/apis/PixverseDto.py
+++ b/comfy_api_nodes/apis/PixverseDto.py
@@ -0,0 +1,57 @@
+# generated by datamodel-codegen:
+#   filename:  filtered-openapi.yaml
+#   timestamp: 2025-04-29T23:44:54+00:00
+
+from __future__ import annotations
+
+from typing import Optional
+
+from pydantic import BaseModel, Field
+
+
+class V2OpenAPII2VResp(BaseModel):
+    video_id: Optional[int] = Field(None, description='Video_id')
+
+
+class V2OpenAPIT2VReq(BaseModel):
+    aspect_ratio: str = Field(
+        ..., description='Aspect ratio (16:9, 4:3, 1:1, 3:4, 9:16)', examples=['16:9']
+    )
+    duration: int = Field(
+        ...,
+        description='Video duration (5, 8 seconds, --model=v3.5 only allows 5,8; --quality=1080p does not support 8s)',
+        examples=[5],
+    )
+    model: str = Field(
+        ..., description='Model version (only supports v3.5)', examples=['v3.5']
+    )
+    motion_mode: Optional[str] = Field(
+        'normal',
+        description='Motion mode (normal, fast, --fast only available when duration=5; --quality=1080p does not support fast)',
+        examples=['normal'],
+    )
+    negative_prompt: Optional[str] = Field(
+        None, description='Negative prompt\n', max_length=2048
+    )
+    prompt: str = Field(..., description='Prompt', max_length=2048)
+    quality: str = Field(
+        ...,
+        description='Video quality ("360p"(Turbo model), "540p", "720p", "1080p")',
+        examples=['540p'],
+    )
+    seed: Optional[int] = Field(None, description='Random seed, range: 0 - 2147483647')
+    style: Optional[str] = Field(
+        None,
+        description='Style (effective when model=v3.5, "anime", "3d_animation", "clay", "comic", "cyberpunk") Do not include style parameter unless needed',
+        examples=['anime'],
+    )
+    template_id: Optional[int] = Field(
+        None,
+        description='Template ID (template_id must be activated before use)',
+        examples=[302325299692608],
+    )
+    water_mark: Optional[bool] = Field(
+        False,
+        description='Watermark (true: add watermark, false: no watermark)',
+        examples=[False],
+    )
--- a/comfy_api_nodes/apis/bfl_api.py
+++ b/comfy_api_nodes/apis/bfl_api.py
@@ -70,29 +70,6 @@ class BFLFluxProGenerateRequest(BaseModel):
    # )


-class Flux2ProGenerateRequest(BaseModel):
-    prompt: str = Field(...)
-    width: int = Field(1024, description="Must be a multiple of 32.")
-    height: int = Field(768, description="Must be a multiple of 32.")
-    seed: int | None = Field(None)
-    prompt_upsampling: bool | None = Field(None)
-    input_image: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_2: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_3: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_4: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_5: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_6: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_7: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_8: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    input_image_9: str | None = Field(None, description="Base64 encoded image for image-to-image generation")
-    safety_tolerance: int | None = Field(
-        5, description="Tolerance level for input and output moderation. Value 0 being most strict.", ge=0, le=5
-    )
-    output_format: str | None = Field(
-        "png", description="Output format for the generated image. Can be 'jpeg' or 'png'."
-    )
-
-
 class BFLFluxKontextProGenerateRequest(BaseModel):
    prompt: str = Field(..., description='The text prompt for what you wannt to edit.')
    input_image: Optional[str] = Field(None, description='Image to edit in base64 format')
@@ -132,9 +109,8 @@ class BFLFluxProUltraGenerateRequest(BaseModel):


 class BFLFluxProGenerateResponse(BaseModel):
-    id: str = Field(..., description="The unique identifier for the generation task.")
-    polling_url: str = Field(..., description="URL to poll for the generation result.")
-    cost: float | None = Field(None, description="Price in cents")
+    id: str = Field(..., description='The unique identifier for the generation task.')
+    polling_url: str = Field(..., description='URL to poll for the generation result.')


 class BFLStatus(str, Enum):
--- a/comfy_api_nodes/apis/client.py
+++ b/comfy_api_nodes/apis/client.py
@@ -0,0 +1,981 @@
+"""
+API Client Framework for api.comfy.org.
+
+This module provides a flexible framework for making API requests from ComfyUI nodes.
+It supports both synchronous and asynchronous API operations with proper type validation.
+
+Key Components:
+--------------
+1. ApiClient - Handles HTTP requests with authentication and error handling
+2. ApiEndpoint - Defines a single HTTP endpoint with its request/response models
+3. ApiOperation - Executes a single synchronous API operation
+
+Usage Examples:
+--------------
+
+# Example 1: Synchronous API Operation
+# ------------------------------------
+# For a simple API call that returns the result immediately:
+
+# 1. Create the API client
+api_client = ApiClient(
+    base_url="https://api.example.com",
+    auth_token="your_auth_token_here",
+    comfy_api_key="your_comfy_api_key_here",
+    timeout=30.0,
+    verify_ssl=True
+)
+
+# 2. Define the endpoint
+user_info_endpoint = ApiEndpoint(
+    path="/v1/users/me",
+    method=HttpMethod.GET,
+    request_model=EmptyRequest,  # No request body needed
+    response_model=UserProfile,   # Pydantic model for the response
+    query_params=None
+)
+
+# 3. Create the request object
+request = EmptyRequest()
+
+# 4. Create and execute the operation
+operation = ApiOperation(
+    endpoint=user_info_endpoint,
+    request=request
+)
+user_profile = await operation.execute(client=api_client)  # Returns immediately with the result
+
+
+# Example 2: Asynchronous API Operation with Polling
+# -------------------------------------------------
+# For an API that starts a task and requires polling for completion:
+
+# 1. Define the endpoints (initial request and polling)
+generate_image_endpoint = ApiEndpoint(
+    path="/v1/images/generate",
+    method=HttpMethod.POST,
+    request_model=ImageGenerationRequest,
+    response_model=TaskCreatedResponse,
+    query_params=None
+)
+
+check_task_endpoint = ApiEndpoint(
+    path="/v1/tasks/{task_id}",
+    method=HttpMethod.GET,
+    request_model=EmptyRequest,
+    response_model=ImageGenerationResult,
+    query_params=None
+)
+
+# 2. Create the request object
+request = ImageGenerationRequest(
+    prompt="a beautiful sunset over mountains",
+    width=1024,
+    height=1024,
+    num_images=1
+)
+
+# 3. Create and execute the polling operation
+operation = PollingOperation(
+    initial_endpoint=generate_image_endpoint,
+    initial_request=request,
+    poll_endpoint=check_task_endpoint,
+    task_id_field="task_id",
+    status_field="status",
+    completed_statuses=["completed"],
+    failed_statuses=["failed", "error"]
+)
+
+# This will make the initial request and then poll until completion
+result = await operation.execute(client=api_client)  # Returns the final ImageGenerationResult when done
+"""
+
+from __future__ import annotations
+import aiohttp
+import asyncio
+import logging
+import io
+import os
+import socket
+from aiohttp.client_exceptions import ClientError, ClientResponseError
+from typing import Type, Optional, Any, TypeVar, Generic, Callable
+from enum import Enum
+import json
+from urllib.parse import urljoin, urlparse
+from pydantic import BaseModel, Field
+import uuid # For generating unique operation IDs
+
+from server import PromptServer
+from comfy.cli_args import args
+from comfy import utils
+from . import request_logger
+
+T = TypeVar("T", bound=BaseModel)
+R = TypeVar("R", bound=BaseModel)
+P = TypeVar("P", bound=BaseModel)  # For poll response
+
+PROGRESS_BAR_MAX = 100
+
+
+class NetworkError(Exception):
+    """Base exception for network-related errors with diagnostic information."""
+    pass
+
+
+class LocalNetworkError(NetworkError):
+    """Exception raised when local network connectivity issues are detected."""
+    pass
+
+
+class ApiServerError(NetworkError):
+    """Exception raised when the API server is unreachable but internet is working."""
+    pass
+
+
+class EmptyRequest(BaseModel):
+    """Base class for empty request bodies.
+    For GET requests, fields will be sent as query parameters."""
+
+    pass
+
+
+class UploadRequest(BaseModel):
+    file_name: str = Field(..., description="Filename to upload")
+    content_type: Optional[str] = Field(
+        None,
+        description="Mime type of the file. For example: image/png, image/jpeg, video/mp4, etc.",
+    )
+
+
+class UploadResponse(BaseModel):
+    download_url: str = Field(..., description="URL to GET uploaded file")
+    upload_url: str = Field(..., description="URL to PUT file to upload")
+
+
+class HttpMethod(str, Enum):
+    GET = "GET"
+    POST = "POST"
+    PUT = "PUT"
+    DELETE = "DELETE"
+    PATCH = "PATCH"
+
+
+class ApiClient:
+    """
+    Client for making HTTP requests to an API with authentication, error handling, and retry logic.
+    """
+
+    def __init__(
+        self,
+        base_url: str,
+        auth_token: Optional[str] = None,
+        comfy_api_key: Optional[str] = None,
+        timeout: float = 3600.0,
+        verify_ssl: bool = True,
+        max_retries: int = 3,
+        retry_delay: float = 1.0,
+        retry_backoff_factor: float = 2.0,
+        retry_status_codes: Optional[tuple[int, ...]] = None,
+        session: Optional[aiohttp.ClientSession] = None,
+    ):
+        self.base_url = base_url
+        self.auth_token = auth_token
+        self.comfy_api_key = comfy_api_key
+        self.timeout = timeout
+        self.verify_ssl = verify_ssl
+        self.max_retries = max_retries
+        self.retry_delay = retry_delay
+        self.retry_backoff_factor = retry_backoff_factor
+        # Default retry status codes: 408 (Request Timeout), 429 (Too Many Requests),
+        # 500, 502, 503, 504 (Server Errors)
+        self.retry_status_codes = retry_status_codes or (408, 429, 500, 502, 503, 504)
+        self._session: Optional[aiohttp.ClientSession] = session
+        self._owns_session = session is None  # Track if we have to close it
+
+    @staticmethod
+    def _generate_operation_id(path: str) -> str:
+        """Generates a unique operation ID for logging."""
+        return f"{path.strip('/').replace('/', '_')}_{uuid.uuid4().hex[:8]}"
+
+    @staticmethod
+    def _create_json_payload_args(
+        data: Optional[dict[str, Any]] = None,
+        headers: Optional[dict[str, str]] = None,
+    ) -> dict[str, Any]:
+        return {
+            "json": data,
+            "headers": headers,
+        }
+
+    def _create_form_data_args(
+        self,
+        data: dict[str, Any] | None,
+        files: dict[str, Any] | None,
+        headers: Optional[dict[str, str]] = None,
+        multipart_parser: Callable | None = None,
+    ) -> dict[str, Any]:
+        if headers and "Content-Type" in headers:
+            del headers["Content-Type"]
+
+        if multipart_parser and data:
+            data = multipart_parser(data)
+
+        if isinstance(data, aiohttp.FormData):
+            form = data  # If the parser already returned a FormData, pass it through
+        else:
+            form = aiohttp.FormData(default_to_multipart=True)
+            if data:  # regular text fields
+                for k, v in data.items():
+                    if v is None:
+                        continue  # aiohttp fails to serialize "None" values
+                    # aiohttp expects strings or bytes; convert enums etc.
+                    form.add_field(k, str(v) if not isinstance(v, (bytes, bytearray)) else v)
+
+        if files:
+            file_iter = files if isinstance(files, list) else files.items()
+            for field_name, file_obj in file_iter:
+                if file_obj is None:
+                    continue  # aiohttp fails to serialize "None" values
+                # file_obj can be (filename, bytes/io.BytesIO, content_type) tuple
+                if isinstance(file_obj, tuple):
+                    filename, file_value, content_type = self._unpack_tuple(file_obj)
+                else:
+                    file_value = file_obj
+                    filename = getattr(file_obj, "name", field_name)
+                    content_type = "application/octet-stream"
+
+                form.add_field(
+                    name=field_name,
+                    value=file_value,
+                    filename=filename,
+                    content_type=content_type,
+                )
+        return {"data": form, "headers": headers or {}}
+
+    @staticmethod
+    def _create_urlencoded_form_data_args(
+        data: dict[str, Any],
+        headers: Optional[dict[str, str]] = None,
+    ) -> dict[str, Any]:
+        headers = headers or {}
+        headers["Content-Type"] = "application/x-www-form-urlencoded"
+        return {
+            "data": data,
+            "headers": headers,
+        }
+
+    def get_headers(self) -> dict[str, str]:
+        """Get headers for API requests, including authentication if available"""
+        headers = {"Content-Type": "application/json", "Accept": "application/json"}
+
+        if self.auth_token:
+            headers["Authorization"] = f"Bearer {self.auth_token}"
+        elif self.comfy_api_key:
+            headers["X-API-KEY"] = self.comfy_api_key
+
+        return headers
+
+    async def _check_connectivity(self, target_url: str) -> dict[str, bool]:
+        """
+        Check connectivity to determine if network issues are local or server-related.
+
+        Args:
+            target_url: URL to check connectivity to
+
+        Returns:
+            Dictionary with connectivity status details
+        """
+        results = {
+            "internet_accessible": False,
+            "api_accessible": False,
+            "is_local_issue": False,
+            "is_api_issue": False,
+        }
+        timeout = aiohttp.ClientTimeout(total=5.0)
+        async with aiohttp.ClientSession(timeout=timeout) as session:
+            try:
+                async with session.get("https://www.google.com", ssl=self.verify_ssl) as resp:
+                    results["internet_accessible"] = resp.status < 500
+            except (ClientError, asyncio.TimeoutError, socket.gaierror):
+                results["is_local_issue"] = True
+                return results  # cannot reach the internet – early exit
+
+            # Now check API health endpoint
+            parsed = urlparse(target_url)
+            health_url = f"{parsed.scheme}://{parsed.netloc}/health"
+            try:
+                async with session.get(health_url, ssl=self.verify_ssl) as resp:
+                    results["api_accessible"] = resp.status < 500
+            except ClientError:
+                pass  # leave as False
+
+        results["is_api_issue"] = results["internet_accessible"] and not results["api_accessible"]
+        return results
+
+    async def request(
+        self,
+        method: str,
+        path: str,
+        params: Optional[dict[str, Any]] = None,
+        data: Optional[dict[str, Any]] = None,
+        files: Optional[dict[str, Any] | list[tuple[str, Any]]] = None,
+        headers: Optional[dict[str, str]] = None,
+        content_type: str = "application/json",
+        multipart_parser: Callable | None = None,
+        retry_count: int = 0,  # Used internally for tracking retries
+    ) -> dict[str, Any]:
+        """
+        Make an HTTP request to the API with automatic retries for transient errors.
+
+        Args:
+            method: HTTP method (GET, POST, etc.)
+            path: API endpoint path (will be joined with base_url)
+            params: Query parameters
+            data: body data
+            files: Files to upload
+            headers: Additional headers
+            content_type: Content type of the request. Defaults to application/json.
+            retry_count: Internal parameter for tracking retries, do not set manually
+
+        Returns:
+            Parsed JSON response
+
+        Raises:
+            LocalNetworkError: If local network connectivity issues are detected
+            ApiServerError: If the API server is unreachable but internet is working
+            Exception: For other request failures
+        """
+
+        # Build full URL and merge headers
+        relative_path = path.lstrip("/")
+        url = urljoin(self.base_url, relative_path)
+        self._check_auth(self.auth_token, self.comfy_api_key)
+
+        request_headers = self.get_headers()
+        if headers:
+            request_headers.update(headers)
+        if files:
+            request_headers.pop("Content-Type", None)
+        if params:
+            params = {k: v for k, v in params.items() if v is not None}  # aiohttp fails to serialize None values
+
+        logging.debug("[DEBUG] Request Headers: %s", request_headers)
+        logging.debug("[DEBUG] Files: %s", files)
+        logging.debug("[DEBUG] Params: %s", params)
+        logging.debug("[DEBUG] Data: %s", data)
+
+        if content_type == "application/x-www-form-urlencoded":
+            payload_args = self._create_urlencoded_form_data_args(data or {}, request_headers)
+        elif content_type == "multipart/form-data":
+            payload_args = self._create_form_data_args(data, files, request_headers, multipart_parser)
+        else:
+            payload_args = self._create_json_payload_args(data, request_headers)
+
+        operation_id = self._generate_operation_id(path)
+        request_logger.log_request_response(
+            operation_id=operation_id,
+            request_method=method,
+            request_url=url,
+            request_headers=request_headers,
+            request_params=params,
+            request_data=data if content_type == "application/json" else "[form-data or other]",
+        )
+
+        session = await self._get_session()
+        try:
+            async with session.request(
+                method,
+                url,
+                params=params,
+                ssl=self.verify_ssl,
+                **payload_args,
+            ) as resp:
+                if resp.status >= 400:
+                    try:
+                        error_data = await resp.json()
+                    except (aiohttp.ContentTypeError, json.JSONDecodeError):
+                        error_data = await resp.text()
+
+                    return await self._handle_http_error(
+                        ClientResponseError(resp.request_info, resp.history, status=resp.status, message=error_data),
+                        operation_id,
+                        method,
+                        url,
+                        params,
+                        data,
+                        files,
+                        headers,
+                        content_type,
+                        multipart_parser,
+                        retry_count=retry_count,
+                        response_content=error_data,
+                    )
+
+                # Success – parse JSON (safely) and log
+                try:
+                    payload = await resp.json()
+                    response_content_to_log = payload
+                except (aiohttp.ContentTypeError, json.JSONDecodeError):
+                    payload = {}
+                    response_content_to_log = await resp.text()
+
+                request_logger.log_request_response(
+                    operation_id=operation_id,
+                    request_method=method,
+                    request_url=url,
+                    response_status_code=resp.status,
+                    response_headers=dict(resp.headers),
+                    response_content=response_content_to_log,
+                )
+                return payload
+
+        except (ClientError, asyncio.TimeoutError, socket.gaierror) as e:
+            # Treat as *connection* problem – optionally retry, else escalate
+            if retry_count < self.max_retries:
+                delay = self.retry_delay * (self.retry_backoff_factor ** retry_count)
+                logging.warning("Connection error. Retrying in %.2fs (%s/%s): %s", delay, retry_count + 1,
+                                self.max_retries, str(e))
+                await asyncio.sleep(delay)
+                return await self.request(
+                    method,
+                    path,
+                    params=params,
+                    data=data,
+                    files=files,
+                    headers=headers,
+                    content_type=content_type,
+                    multipart_parser=multipart_parser,
+                    retry_count=retry_count + 1,
+                )
+            # One final connectivity check for diagnostics
+            connectivity = await self._check_connectivity(self.base_url)
+            if connectivity["is_local_issue"]:
+                raise LocalNetworkError(
+                    "Unable to connect to the API server due to local network issues. "
+                    "Please check your internet connection and try again."
+                ) from e
+            raise ApiServerError(
+                f"The API server at {self.base_url} is currently unreachable. "
+                f"The service may be experiencing issues. Please try again later."
+            ) from e
+
+    @staticmethod
+    def _check_auth(auth_token, comfy_api_key):
+        """Verify that an auth token is present or comfy_api_key is present"""
+        if auth_token is None and comfy_api_key is None:
+            raise Exception("Unauthorized: Please login first to use this node.")
+        return auth_token or comfy_api_key
+
+    @staticmethod
+    async def upload_file(
+        upload_url: str,
+        file: io.BytesIO | str,
+        content_type: str | None = None,
+        max_retries: int = 3,
+        retry_delay: float = 1.0,
+        retry_backoff_factor: float = 2.0,
+    ) -> aiohttp.ClientResponse:
+        """Upload a file to the API with retry logic.
+
+        Args:
+            upload_url: The URL to upload to
+            file: Either a file path string, BytesIO object, or tuple of (file_path, filename)
+            content_type: Optional mime type to set for the upload
+            max_retries: Maximum number of retry attempts
+            retry_delay: Initial delay between retries in seconds
+            retry_backoff_factor: Multiplier for the delay after each retry
+        """
+        headers: dict[str, str] = {}
+        skip_auto_headers: set[str] = set()
+        if content_type:
+            headers["Content-Type"] = content_type
+        else:
+            # tell aiohttp not to add Content-Type that will break the request signature and result in a 403 status.
+            skip_auto_headers.add("Content-Type")
+
+        # Extract file bytes
+        if isinstance(file, io.BytesIO):
+            file.seek(0)
+            data = file.read()
+        elif isinstance(file, str):
+            with open(file, "rb") as f:
+                data = f.read()
+        else:
+            raise ValueError("File must be BytesIO or str path")
+
+        parsed = urlparse(upload_url)
+        basename = os.path.basename(parsed.path) or parsed.netloc or "upload"
+        operation_id = f"upload_{basename}_{uuid.uuid4().hex[:8]}"
+        request_logger.log_request_response(
+            operation_id=operation_id,
+            request_method="PUT",
+            request_url=upload_url,
+            request_headers=headers,
+            request_data=f"[File data {len(data)} bytes]",
+        )
+
+        delay = retry_delay
+        for attempt in range(max_retries + 1):
+            try:
+                timeout = aiohttp.ClientTimeout(total=None)  # honour server side timeouts
+                async with aiohttp.ClientSession(timeout=timeout) as session:
+                    async with session.put(
+                        upload_url, data=data, headers=headers, skip_auto_headers=skip_auto_headers,
+                    ) as resp:
+                        resp.raise_for_status()
+                        request_logger.log_request_response(
+                            operation_id=operation_id,
+                            request_method="PUT",
+                            request_url=upload_url,
+                            response_status_code=resp.status,
+                            response_headers=dict(resp.headers),
+                            response_content="File uploaded successfully.",
+                        )
+                        return resp
+            except (ClientError, asyncio.TimeoutError) as e:
+                request_logger.log_request_response(
+                    operation_id=operation_id,
+                    request_method="PUT",
+                    request_url=upload_url,
+                    response_status_code=e.status if hasattr(e, "status") else None,
+                    response_headers=dict(e.headers) if hasattr(e, "headers") else None,
+                    response_content=None,
+                    error_message=f"{type(e).__name__}: {str(e)}",
+                )
+                if attempt < max_retries:
+                    logging.warning(
+                        "Upload failed (%s/%s). Retrying in %.2fs. %s", attempt + 1, max_retries, delay, str(e)
+                    )
+                    await asyncio.sleep(delay)
+                    delay *= retry_backoff_factor
+                else:
+                    raise NetworkError(f"Failed to upload file after {max_retries + 1} attempts: {e}") from e
+
+    async def _handle_http_error(
+        self,
+        exc: ClientResponseError,
+        operation_id: str,
+        *req_meta,
+        retry_count: int,
+        response_content: dict | str = "",
+    ) -> dict[str, Any]:
+        status_code = exc.status
+        if status_code == 401:
+            user_friendly = "Unauthorized: Please login first to use this node."
+        elif status_code == 402:
+            user_friendly = "Payment Required: Please add credits to your account to use this node."
+        elif status_code == 409:
+            user_friendly = "There is a problem with your account. Please contact support@comfy.org."
+        elif status_code == 429:
+            user_friendly = "Rate Limit Exceeded: Please try again later."
+        else:
+            if isinstance(response_content, dict):
+                if "error" in response_content and "message" in response_content["error"]:
+                    user_friendly = f"API Error: {response_content['error']['message']}"
+                    if "type" in response_content["error"]:
+                        user_friendly += f" (Type: {response_content['error']['type']})"
+                else: # Handle cases where error is just a JSON dict with unknown format
+                    user_friendly = f"API Error: {json.dumps(response_content)}"
+            else:
+                if len(response_content) < 200:  # Arbitrary limit for display
+                    user_friendly = f"API Error (raw): {response_content}"
+                else:
+                    user_friendly = f"API Error (raw, status {response_content})"
+
+        request_logger.log_request_response(
+            operation_id=operation_id,
+            request_method=req_meta[0],
+            request_url=req_meta[1],
+            response_status_code=exc.status,
+            response_headers=dict(req_meta[5]) if req_meta[5] else None,
+            response_content=response_content,
+            error_message=f"HTTP Error {exc.status}",
+        )
+
+        logging.debug("[DEBUG] API Error: %s (Status: %s)", user_friendly, status_code)
+        if response_content:
+            logging.debug("[DEBUG] Response content: %s", response_content)
+
+        # Retry if eligible
+        if status_code in self.retry_status_codes and retry_count < self.max_retries:
+            delay = self.retry_delay * (self.retry_backoff_factor ** retry_count)
+            logging.warning(
+                "HTTP error %s. Retrying in %.2fs (%s/%s)",
+                status_code,
+                delay,
+                retry_count + 1,
+                self.max_retries,
+            )
+            await asyncio.sleep(delay)
+            return await self.request(
+                req_meta[0],  # method
+                req_meta[1].replace(self.base_url, ""),  # path
+                params=req_meta[2],
+                data=req_meta[3],
+                files=req_meta[4],
+                headers=req_meta[5],
+                content_type=req_meta[6],
+                multipart_parser=req_meta[7],
+                retry_count=retry_count + 1,
+            )
+
+        raise Exception(user_friendly) from exc
+
+    @staticmethod
+    def _unpack_tuple(t):
+        """Helper to normalise (filename, file, content_type) tuples."""
+        if len(t) == 3:
+            return t
+        elif len(t) == 2:
+            return t[0], t[1], "application/octet-stream"
+        else:
+            raise ValueError("files tuple must be (filename, file[, content_type])")
+
+    async def _get_session(self) -> aiohttp.ClientSession:
+        if self._session is None or self._session.closed:
+            timeout = aiohttp.ClientTimeout(total=self.timeout)
+            self._session = aiohttp.ClientSession(timeout=timeout)
+            self._owns_session = True
+        return self._session
+
+    async def close(self) -> None:
+        if self._owns_session and self._session and not self._session.closed:
+            await self._session.close()
+
+    async def __aenter__(self) -> "ApiClient":
+        """Allow usage as async‑context‑manager – ensures clean teardown"""
+        return self
+
+    async def __aexit__(self, exc_type, exc, tb):
+        await self.close()
+
+
+class ApiEndpoint(Generic[T, R]):
+    """Defines an API endpoint with its request and response types"""
+
+    def __init__(
+        self,
+        path: str,
+        method: HttpMethod,
+        request_model: Type[T],
+        response_model: Type[R],
+        query_params: Optional[dict[str, Any]] = None,
+    ):
+        """Initialize an API endpoint definition.
+
+        Args:
+            path: The URL path for this endpoint, can include placeholders like {id}
+            method: The HTTP method to use (GET, POST, etc.)
+            request_model: Pydantic model class that defines the structure and validation rules for API requests to this endpoint
+            response_model: Pydantic model class that defines the structure and validation rules for API responses from this endpoint
+            query_params: Optional dictionary of query parameters to include in the request
+        """
+        self.path = path
+        self.method = method
+        self.request_model = request_model
+        self.response_model = response_model
+        self.query_params = query_params or {}
+
+
+class SynchronousOperation(Generic[T, R]):
+    """Represents a single synchronous API operation."""
+
+    def __init__(
+        self,
+        endpoint: ApiEndpoint[T, R],
+        request: T,
+        files: Optional[dict[str, Any] | list[tuple[str, Any]]] = None,
+        api_base: str | None = None,
+        auth_token: Optional[str] = None,
+        comfy_api_key: Optional[str] = None,
+        auth_kwargs: Optional[dict[str, str]] = None,
+        timeout: float = 7200.0,
+        verify_ssl: bool = True,
+        content_type: str = "application/json",
+        multipart_parser: Callable | None = None,
+        max_retries: int = 3,
+        retry_delay: float = 1.0,
+        retry_backoff_factor: float = 2.0,
+    ) -> None:
+        self.endpoint = endpoint
+        self.request = request
+        self.files = files
+        self.api_base: str = api_base or args.comfy_api_base
+        self.auth_token = auth_token
+        self.comfy_api_key = comfy_api_key
+        if auth_kwargs is not None:
+            self.auth_token = auth_kwargs.get("auth_token", self.auth_token)
+            self.comfy_api_key = auth_kwargs.get("comfy_api_key", self.comfy_api_key)
+        self.timeout = timeout
+        self.verify_ssl = verify_ssl
+        self.content_type = content_type
+        self.multipart_parser = multipart_parser
+        self.max_retries = max_retries
+        self.retry_delay = retry_delay
+        self.retry_backoff_factor = retry_backoff_factor
+
+    async def execute(self, client: Optional[ApiClient] = None) -> R:
+        owns_client = client is None
+        if owns_client:
+            client = ApiClient(
+                base_url=self.api_base,
+                auth_token=self.auth_token,
+                comfy_api_key=self.comfy_api_key,
+                timeout=self.timeout,
+                verify_ssl=self.verify_ssl,
+                max_retries=self.max_retries,
+                retry_delay=self.retry_delay,
+                retry_backoff_factor=self.retry_backoff_factor,
+            )
+
+        try:
+            request_dict: Optional[dict[str, Any]]
+            if isinstance(self.request, EmptyRequest):
+                request_dict = None
+            else:
+                request_dict = self.request.model_dump(exclude_none=True)
+                for k, v in list(request_dict.items()):
+                    if isinstance(v, Enum):
+                        request_dict[k] = v.value
+
+            logging.debug("[DEBUG] API Request: %s %s", self.endpoint.method.value, self.endpoint.path)
+            logging.debug("[DEBUG] Request Data: %s", json.dumps(request_dict, indent=2))
+            logging.debug("[DEBUG] Query Params: %s", self.endpoint.query_params)
+
+            response_json = await client.request(
+                self.endpoint.method.value,
+                self.endpoint.path,
+                params=self.endpoint.query_params,
+                data=request_dict,
+                files=self.files,
+                content_type=self.content_type,
+                multipart_parser=self.multipart_parser,
+            )
+
+            logging.debug("=" * 50)
+            logging.debug("[DEBUG] RESPONSE DETAILS:")
+            logging.debug("[DEBUG] Status Code: 200 (Success)")
+            logging.debug("[DEBUG] Response Body: %s", json.dumps(response_json, indent=2))
+            logging.debug("=" * 50)
+
+            parsed_response = self.endpoint.response_model.model_validate(response_json)
+            logging.debug("[DEBUG] Parsed Response: %s", parsed_response)
+            return parsed_response
+        finally:
+            if owns_client:
+                await client.close()
+
+
+class TaskStatus(str, Enum):
+    """Enum for task status values"""
+
+    COMPLETED = "completed"
+    FAILED = "failed"
+    PENDING = "pending"
+
+
+class PollingOperation(Generic[T, R]):
+    """Represents an asynchronous API operation that requires polling for completion."""
+
+    def __init__(
+        self,
+        poll_endpoint: ApiEndpoint[EmptyRequest, R],
+        completed_statuses: list[str],
+        failed_statuses: list[str],
+        *,
+        status_extractor: Callable[[R], Optional[str]],
+        progress_extractor: Callable[[R], Optional[float]] | None = None,
+        result_url_extractor: Callable[[R], Optional[str]] | None = None,
+        price_extractor: Callable[[R], Optional[float]] | None = None,
+        request: Optional[T] = None,
+        api_base: str | None = None,
+        auth_token: Optional[str] = None,
+        comfy_api_key: Optional[str] = None,
+        auth_kwargs: Optional[dict[str, str]] = None,
+        poll_interval: float = 5.0,
+        max_poll_attempts: int = 120,  # Default max polling attempts (10 minutes with 5s interval)
+        max_retries: int = 3,  # Max retries per individual API call
+        retry_delay: float = 1.0,
+        retry_backoff_factor: float = 2.0,
+        estimated_duration: Optional[float] = None,
+        node_id: Optional[str] = None,
+    ) -> None:
+        self.poll_endpoint = poll_endpoint
+        self.request = request
+        self.api_base: str = api_base or args.comfy_api_base
+        self.auth_token = auth_token
+        self.comfy_api_key = comfy_api_key
+        if auth_kwargs is not None:
+            self.auth_token = auth_kwargs.get("auth_token", self.auth_token)
+            self.comfy_api_key = auth_kwargs.get("comfy_api_key", self.comfy_api_key)
+        self.poll_interval = poll_interval
+        self.max_poll_attempts = max_poll_attempts
+        self.max_retries = max_retries
+        self.retry_delay = retry_delay
+        self.retry_backoff_factor = retry_backoff_factor
+        self.estimated_duration = estimated_duration
+        self.status_extractor = status_extractor or (lambda x: getattr(x, "status", None))
+        self.progress_extractor = progress_extractor
+        self.result_url_extractor = result_url_extractor
+        self.price_extractor = price_extractor
+        self.node_id = node_id
+        self.completed_statuses = completed_statuses
+        self.failed_statuses = failed_statuses
+        self.final_response: Optional[R] = None
+        self.extracted_price: Optional[float] = None
+
+    async def execute(self, client: Optional[ApiClient] = None) -> R:
+        owns_client = client is None
+        if owns_client:
+            client = ApiClient(
+                base_url=self.api_base,
+                auth_token=self.auth_token,
+                comfy_api_key=self.comfy_api_key,
+                max_retries=self.max_retries,
+                retry_delay=self.retry_delay,
+                retry_backoff_factor=self.retry_backoff_factor,
+            )
+        try:
+            return await self._poll_until_complete(client)
+        finally:
+            if owns_client:
+                await client.close()
+
+    def _display_text_on_node(self, text: str):
+        if not self.node_id:
+            return
+        if self.extracted_price is not None:
+            text = f"Price: ${self.extracted_price}\n{text}"
+        PromptServer.instance.send_progress_text(text, self.node_id)
+
+    def _display_time_progress_on_node(self, time_completed: int | float):
+        if not self.node_id:
+            return
+        if self.estimated_duration is not None:
+            remaining = max(0, int(self.estimated_duration) - time_completed)
+            message = f"Task in progress: {time_completed}s (~{remaining}s remaining)"
+        else:
+            message = f"Task in progress: {time_completed}s"
+        self._display_text_on_node(message)
+
+    def _check_task_status(self, response: R) -> TaskStatus:
+        try:
+            status = self.status_extractor(response)
+            if status in self.completed_statuses:
+                return TaskStatus.COMPLETED
+            if status in self.failed_statuses:
+                return TaskStatus.FAILED
+            return TaskStatus.PENDING
+        except Exception as e:
+            logging.error("Error extracting status: %s", e)
+            return TaskStatus.PENDING
+
+    async def _poll_until_complete(self, client: ApiClient) -> R:
+        """Poll until the task is complete"""
+        consecutive_errors = 0
+        max_consecutive_errors = min(5, self.max_retries * 2)  # Limit consecutive errors
+
+        if self.progress_extractor:
+            progress = utils.ProgressBar(PROGRESS_BAR_MAX)
+
+        status = TaskStatus.PENDING
+        for poll_count in range(1, self.max_poll_attempts + 1):
+            try:
+                logging.debug("[DEBUG] Polling attempt #%s", poll_count)
+
+                request_dict = None if self.request is None else self.request.model_dump(exclude_none=True)
+
+                if poll_count == 1:
+                    logging.debug(
+                        "[DEBUG] Poll Request: %s %s",
+                        self.poll_endpoint.method.value,
+                        self.poll_endpoint.path,
+                    )
+                    logging.debug(
+                        "[DEBUG] Poll Request Data: %s",
+                        json.dumps(request_dict, indent=2) if request_dict else "None",
+                    )
+
+                # Query task status
+                resp = await client.request(
+                    self.poll_endpoint.method.value,
+                    self.poll_endpoint.path,
+                    params=self.poll_endpoint.query_params,
+                    data=request_dict,
+                )
+                consecutive_errors = 0  # reset on success
+                response_obj: R = self.poll_endpoint.response_model.model_validate(resp)
+
+                # Check if task is complete
+                status = self._check_task_status(response_obj)
+                logging.debug("[DEBUG] Task Status: %s", status)
+
+                # If progress extractor is provided, extract progress
+                if self.progress_extractor:
+                    new_progress = self.progress_extractor(response_obj)
+                    if new_progress is not None:
+                        progress.update_absolute(new_progress, total=PROGRESS_BAR_MAX)
+
+                if self.price_extractor:
+                    price = self.price_extractor(response_obj)
+                    if price is not None:
+                        self.extracted_price = price
+
+                if status == TaskStatus.COMPLETED:
+                    message = "Task completed successfully"
+                    if self.result_url_extractor:
+                        result_url = self.result_url_extractor(response_obj)
+                        if result_url:
+                            message = f"Result URL: {result_url}"
+                    logging.debug("[DEBUG] %s", message)
+                    self._display_text_on_node(message)
+                    self.final_response = response_obj
+                    if self.progress_extractor:
+                        progress.update(100)
+                    return self.final_response
+                if status == TaskStatus.FAILED:
+                    message = f"Task failed: {json.dumps(resp)}"
+                    logging.error("[DEBUG] %s", message)
+                    raise Exception(message)
+                logging.debug("[DEBUG] Task still pending, continuing to poll...")
+                # Task pending – wait
+                for i in range(int(self.poll_interval)):
+                    self._display_time_progress_on_node((poll_count - 1) * self.poll_interval + i)
+                    await asyncio.sleep(1)
+
+            except (LocalNetworkError, ApiServerError, NetworkError) as e:
+                consecutive_errors += 1
+                if consecutive_errors >= max_consecutive_errors:
+                    raise Exception(
+                        f"Polling aborted after {consecutive_errors} network errors: {str(e)}"
+                    ) from e
+                logging.warning(
+                    "Network error (%s/%s): %s",
+                    consecutive_errors,
+                    max_consecutive_errors,
+                    str(e),
+                )
+                await asyncio.sleep(self.poll_interval)
+            except Exception as e:
+                # For other errors, increment count and potentially abort
+                consecutive_errors += 1
+                if consecutive_errors >= max_consecutive_errors or status == TaskStatus.FAILED:
+                    raise Exception(
+                        f"Polling aborted after {consecutive_errors} consecutive errors: {str(e)}"
+                    ) from e
+
+                logging.error("[DEBUG] Polling error: %s", str(e))
+                logging.warning(
+                    "Error during polling (attempt %s/%s): %s. Will retry in %s seconds.",
+                    poll_count,
+                    self.max_poll_attempts,
+                    str(e),
+                    self.poll_interval,
+                )
+                await asyncio.sleep(self.poll_interval)
+
+        # If we've exhausted all polling attempts
+        raise Exception(
+            f"Polling timed out after {self.max_poll_attempts} attempts (" f"{self.max_poll_attempts * self.poll_interval} seconds). "
+            "The operation may still be running on the server but is taking longer than expected."
+        )
--- a/comfy_api_nodes/apis/gemini_api.py
+++ b/comfy_api_nodes/apis/gemini_api.py
@@ -1,236 +1,22 @@
-from datetime import date
-from enum import Enum
-from typing import Any
+from typing import Optional

-from pydantic import BaseModel, Field
-
-
-class GeminiSafetyCategory(str, Enum):
-    HARM_CATEGORY_SEXUALLY_EXPLICIT = "HARM_CATEGORY_SEXUALLY_EXPLICIT"
-    HARM_CATEGORY_HATE_SPEECH = "HARM_CATEGORY_HATE_SPEECH"
-    HARM_CATEGORY_HARASSMENT = "HARM_CATEGORY_HARASSMENT"
-    HARM_CATEGORY_DANGEROUS_CONTENT = "HARM_CATEGORY_DANGEROUS_CONTENT"
-
-
-class GeminiSafetyThreshold(str, Enum):
-    OFF = "OFF"
-    BLOCK_NONE = "BLOCK_NONE"
-    BLOCK_LOW_AND_ABOVE = "BLOCK_LOW_AND_ABOVE"
-    BLOCK_MEDIUM_AND_ABOVE = "BLOCK_MEDIUM_AND_ABOVE"
-    BLOCK_ONLY_HIGH = "BLOCK_ONLY_HIGH"
-
-
-class GeminiSafetySetting(BaseModel):
-    category: GeminiSafetyCategory
-    threshold: GeminiSafetyThreshold
-
-
-class GeminiRole(str, Enum):
-    user = "user"
-    model = "model"
-
-
-class GeminiMimeType(str, Enum):
-    application_pdf = "application/pdf"
-    audio_mpeg = "audio/mpeg"
-    audio_mp3 = "audio/mp3"
-    audio_wav = "audio/wav"
-    image_png = "image/png"
-    image_jpeg = "image/jpeg"
-    image_webp = "image/webp"
-    text_plain = "text/plain"
-    video_mov = "video/mov"
-    video_mpeg = "video/mpeg"
-    video_mp4 = "video/mp4"
-    video_mpg = "video/mpg"
-    video_avi = "video/avi"
-    video_wmv = "video/wmv"
-    video_mpegps = "video/mpegps"
-    video_flv = "video/flv"
-
-
-class GeminiInlineData(BaseModel):
-    data: str | None = Field(
-        None,
-        description="The base64 encoding of the image, PDF, or video to include inline in the prompt. "
-        "When including media inline, you must also specify the media type (mimeType) of the data. Size limit: 20MB",
-    )
-    mimeType: GeminiMimeType | None = Field(None)
-
-
-class GeminiFileData(BaseModel):
-    fileUri: str | None = Field(None)
-    mimeType: GeminiMimeType | None = Field(None)
-
-
-class GeminiPart(BaseModel):
-    inlineData: GeminiInlineData | None = Field(None)
-    fileData: GeminiFileData | None = Field(None)
-    text: str | None = Field(None)
-
-
-class GeminiTextPart(BaseModel):
-    text: str | None = Field(None)
-
-
-class GeminiContent(BaseModel):
-    parts: list[GeminiPart] = Field([])
-    role: GeminiRole = Field(..., examples=["user"])
-
-
-class GeminiSystemInstructionContent(BaseModel):
-    parts: list[GeminiTextPart] = Field(
-        ...,
-        description="A list of ordered parts that make up a single message. "
-        "Different parts may have different IANA MIME types.",
-    )
-    role: GeminiRole = Field(
-        ...,
-        description="The identity of the entity that creates the message. "
-        "The following values are supported: "
-        "user: This indicates that the message is sent by a real person, typically a user-generated message. "
-        "model: This indicates that the message is generated by the model. "
-        "The model value is used to insert messages from model into the conversation during multi-turn conversations. "
-        "For non-multi-turn conversations, this field can be left blank or unset.",
-    )
-
-
-class GeminiFunctionDeclaration(BaseModel):
-    description: str | None = Field(None)
-    name: str = Field(...)
-    parameters: dict[str, Any] = Field(..., description="JSON schema for the function parameters")
-
-
-class GeminiTool(BaseModel):
-    functionDeclarations: list[GeminiFunctionDeclaration] | None = Field(None)
-
-
-class GeminiOffset(BaseModel):
-    nanos: int | None = Field(None, ge=0, le=999999999)
-    seconds: int | None = Field(None, ge=-315576000000, le=315576000000)
-
-
-class GeminiVideoMetadata(BaseModel):
-    endOffset: GeminiOffset | None = Field(None)
-    startOffset: GeminiOffset | None = Field(None)
-
-
-class GeminiGenerationConfig(BaseModel):
-    maxOutputTokens: int | None = Field(None, ge=16, le=8192)
-    seed: int | None = Field(None)
-    stopSequences: list[str] | None = Field(None)
-    temperature: float | None = Field(None, ge=0.0, le=2.0)
-    topK: int | None = Field(None, ge=1)
-    topP: float | None = Field(None, ge=0.0, le=1.0)
+from comfy_api_nodes.apis import GeminiGenerationConfig, GeminiContent, GeminiSafetySetting, GeminiSystemInstructionContent, GeminiTool, GeminiVideoMetadata
+from pydantic import BaseModel


 class GeminiImageConfig(BaseModel):
-    aspectRatio: str | None = Field(None)
-    imageSize: str | None = Field(None)
+    aspectRatio: Optional[str] = None


 class GeminiImageGenerationConfig(GeminiGenerationConfig):
-    responseModalities: list[str] | None = Field(None)
-    imageConfig: GeminiImageConfig | None = Field(None)
+    responseModalities: Optional[list[str]] = None
+    imageConfig: Optional[GeminiImageConfig] = None


 class GeminiImageGenerateContentRequest(BaseModel):
-    contents: list[GeminiContent] = Field(...)
-    generationConfig: GeminiImageGenerationConfig | None = Field(None)
-    safetySettings: list[GeminiSafetySetting] | None = Field(None)
-    systemInstruction: GeminiSystemInstructionContent | None = Field(None)
-    tools: list[GeminiTool] | None = Field(None)
-    videoMetadata: GeminiVideoMetadata | None = Field(None)
-
-
-class GeminiGenerateContentRequest(BaseModel):
-    contents: list[GeminiContent] = Field(...)
-    generationConfig: GeminiGenerationConfig | None = Field(None)
-    safetySettings: list[GeminiSafetySetting] | None = Field(None)
-    systemInstruction: GeminiSystemInstructionContent | None = Field(None)
-    tools: list[GeminiTool] | None = Field(None)
-    videoMetadata: GeminiVideoMetadata | None = Field(None)
-
-
-class Modality(str, Enum):
-    MODALITY_UNSPECIFIED = "MODALITY_UNSPECIFIED"
-    TEXT = "TEXT"
-    IMAGE = "IMAGE"
-    VIDEO = "VIDEO"
-    AUDIO = "AUDIO"
-    DOCUMENT = "DOCUMENT"
-
-
-class ModalityTokenCount(BaseModel):
-    modality: Modality | None = None
-    tokenCount: int | None = Field(None, description="Number of tokens for the given modality.")
-
-
-class Probability(str, Enum):
-    NEGLIGIBLE = "NEGLIGIBLE"
-    LOW = "LOW"
-    MEDIUM = "MEDIUM"
-    HIGH = "HIGH"
-    UNKNOWN = "UNKNOWN"
-
-
-class GeminiSafetyRating(BaseModel):
-    category: GeminiSafetyCategory | None = None
-    probability: Probability | None = Field(
-        None,
-        description="The probability that the content violates the specified safety category",
-    )
-
-
-class GeminiCitation(BaseModel):
-    authors: list[str] | None = None
-    endIndex: int | None = None
-    license: str | None = None
-    publicationDate: date | None = None
-    startIndex: int | None = None
-    title: str | None = None
-    uri: str | None = None
-
-
-class GeminiCitationMetadata(BaseModel):
-    citations: list[GeminiCitation] | None = None
-
-
-class GeminiCandidate(BaseModel):
-    citationMetadata: GeminiCitationMetadata | None = None
-    content: GeminiContent | None = None
-    finishReason: str | None = None
-    safetyRatings: list[GeminiSafetyRating] | None = None
-
-
-class GeminiPromptFeedback(BaseModel):
-    blockReason: str | None = None
-    blockReasonMessage: str | None = None
-    safetyRatings: list[GeminiSafetyRating] | None = None
-
-
-class GeminiUsageMetadata(BaseModel):
-    cachedContentTokenCount: int | None = Field(
-        None,
-        description="Output only. Number of tokens in the cached part in the input (the cached content).",
-    )
-    candidatesTokenCount: int | None = Field(None, description="Number of tokens in the response(s).")
-    candidatesTokensDetails: list[ModalityTokenCount] | None = Field(
-        None, description="Breakdown of candidate tokens by modality."
-    )
-    promptTokenCount: int | None = Field(
-        None,
-        description="Number of tokens in the request. When cachedContent is set, this is still the total effective prompt size meaning this includes the number of tokens in the cached content.",
-    )
-    promptTokensDetails: list[ModalityTokenCount] | None = Field(
-        None, description="Breakdown of prompt tokens by modality."
-    )
-    thoughtsTokenCount: int | None = Field(None, description="Number of tokens present in thoughts output.")
-    toolUsePromptTokenCount: int | None = Field(None, description="Number of tokens present in tool-use prompt(s).")
-
-
-class GeminiGenerateContentResponse(BaseModel):
-    candidates: list[GeminiCandidate] | None = Field(None)
-    promptFeedback: GeminiPromptFeedback | None = Field(None)
-    usageMetadata: GeminiUsageMetadata | None = Field(None)
-    modelVersion: str | None = Field(None)
+    contents: list[GeminiContent]
+    generationConfig: Optional[GeminiImageGenerationConfig] = None
+    safetySettings: Optional[list[GeminiSafetySetting]] = None
+    systemInstruction: Optional[GeminiSystemInstructionContent] = None
+    tools: Optional[list[GeminiTool]] = None
+    videoMetadata: Optional[GeminiVideoMetadata] = None
--- a/comfy_api_nodes/apis/kling_api.py
+++ b/comfy_api_nodes/apis/kling_api.py
@@ -1,66 +0,0 @@
-from pydantic import BaseModel, Field
-
-
-class OmniProText2VideoRequest(BaseModel):
-    model_name: str = Field(..., description="kling-video-o1")
-    aspect_ratio: str = Field(..., description="'16:9', '9:16' or '1:1'")
-    duration: str = Field(..., description="'5' or '10'")
-    prompt: str = Field(...)
-    mode: str = Field("pro")
-
-
-class OmniParamImage(BaseModel):
-    image_url: str = Field(...)
-    type: str | None = Field(None, description="Can be 'first_frame' or 'end_frame'")
-
-
-class OmniParamVideo(BaseModel):
-    video_url: str = Field(...)
-    refer_type: str | None = Field(..., description="Can be 'base' or 'feature'")
-    keep_original_sound: str = Field(..., description="'yes' or 'no'")
-
-
-class OmniProFirstLastFrameRequest(BaseModel):
-    model_name: str = Field(..., description="kling-video-o1")
-    image_list: list[OmniParamImage] = Field(..., min_length=1, max_length=7)
-    duration: str = Field(..., description="'5' or '10'")
-    prompt: str = Field(...)
-    mode: str = Field("pro")
-
-
-class OmniProReferences2VideoRequest(BaseModel):
-    model_name: str = Field(..., description="kling-video-o1")
-    aspect_ratio: str | None = Field(..., description="'16:9', '9:16' or '1:1'")
-    image_list: list[OmniParamImage] | None = Field(
-        None, max_length=7, description="Max length 4 when video is present."
-    )
-    video_list: list[OmniParamVideo] | None = Field(None, max_length=1)
-    duration: str | None = Field(..., description="From 3 to 10.")
-    prompt: str = Field(...)
-    mode: str = Field("pro")
-
-
-class TaskStatusVideoResult(BaseModel):
-    duration: str | None = Field(None, description="Total video duration")
-    id: str | None = Field(None, description="Generated video ID")
-    url: str | None = Field(None, description="URL for generated video")
-
-
-class TaskStatusVideoResults(BaseModel):
-    videos: list[TaskStatusVideoResult] | None = Field(None)
-
-
-class TaskStatusVideoResponseData(BaseModel):
-    created_at: int | None = Field(None, description="Task creation time")
-    updated_at: int | None = Field(None, description="Task update time")
-    task_status: str | None = None
-    task_status_msg: str | None = Field(None, description="Additional failure reason. Only for polling endpoint.")
-    task_id: str | None = Field(None, description="Task ID")
-    task_result: TaskStatusVideoResults | None = Field(None)
-
-
-class TaskStatusVideoResponse(BaseModel):
-    code: int | None = Field(None, description="Error code")
-    message: str | None = Field(None, description="Error message")
-    request_id: str | None = Field(None, description="Request ID")
-    data: TaskStatusVideoResponseData | None = Field(None)
--- a/comfy_api_nodes/apis/minimax_api.py
+++ b/comfy_api_nodes/apis/minimax_api.py
@@ -1,120 +0,0 @@
-from enum import Enum
-from typing import Optional
-
-from pydantic import BaseModel, Field
-
-
-class MinimaxBaseResponse(BaseModel):
-    status_code: int = Field(
-        ...,
-        description='Status code. 0 indicates success, other values indicate errors.',
-    )
-    status_msg: str = Field(
-        ..., description='Specific error details or success message.'
-    )
-
-
-class File(BaseModel):
-    bytes: Optional[int] = Field(None, description='File size in bytes')
-    created_at: Optional[int] = Field(
-        None, description='Unix timestamp when the file was created, in seconds'
-    )
-    download_url: Optional[str] = Field(
-        None, description='The URL to download the video'
-    )
-    backup_download_url: Optional[str] = Field(
-        None, description='The backup URL to download the video'
-    )
-
-    file_id: Optional[int] = Field(None, description='Unique identifier for the file')
-    filename: Optional[str] = Field(None, description='The name of the file')
-    purpose: Optional[str] = Field(None, description='The purpose of using the file')
-
-
-class MinimaxFileRetrieveResponse(BaseModel):
-    base_resp: MinimaxBaseResponse
-    file: File
-
-
-class MiniMaxModel(str, Enum):
-    T2V_01_Director = 'T2V-01-Director'
-    I2V_01_Director = 'I2V-01-Director'
-    S2V_01 = 'S2V-01'
-    I2V_01 = 'I2V-01'
-    I2V_01_live = 'I2V-01-live'
-    T2V_01 = 'T2V-01'
-    Hailuo_02 = 'MiniMax-Hailuo-02'
-
-
-class Status6(str, Enum):
-    Queueing = 'Queueing'
-    Preparing = 'Preparing'
-    Processing = 'Processing'
-    Success = 'Success'
-    Fail = 'Fail'
-
-
-class MinimaxTaskResultResponse(BaseModel):
-    base_resp: MinimaxBaseResponse
-    file_id: Optional[str] = Field(
-        None,
-        description='After the task status changes to Success, this field returns the file ID corresponding to the generated video.',
-    )
-    status: Status6 = Field(
-        ...,
-        description="Task status: 'Queueing' (in queue), 'Preparing' (task is preparing), 'Processing' (generating), 'Success' (task completed successfully), or 'Fail' (task failed).",
-    )
-    task_id: str = Field(..., description='The task ID being queried.')
-
-
-class SubjectReferenceItem(BaseModel):
-    image: Optional[str] = Field(
-        None, description='URL or base64 encoding of the subject reference image.'
-    )
-    mask: Optional[str] = Field(
-        None,
-        description='URL or base64 encoding of the mask for the subject reference image.',
-    )
-
-
-class MinimaxVideoGenerationRequest(BaseModel):
-    callback_url: Optional[str] = Field(
-        None,
-        description='Optional. URL to receive real-time status updates about the video generation task.',
-    )
-    first_frame_image: Optional[str] = Field(
-        None,
-        description='URL or base64 encoding of the first frame image. Required when model is I2V-01, I2V-01-Director, or I2V-01-live.',
-    )
-    model: MiniMaxModel = Field(
-        ...,
-        description='Required. ID of model. Options: T2V-01-Director, I2V-01-Director, S2V-01, I2V-01, I2V-01-live, T2V-01',
-    )
-    prompt: Optional[str] = Field(
-        None,
-        description='Description of the video. Should be less than 2000 characters. Supports camera movement instructions in [brackets].',
-        max_length=2000,
-    )
-    prompt_optimizer: Optional[bool] = Field(
-        True,
-        description='If true (default), the model will automatically optimize the prompt. Set to false for more precise control.',
-    )
-    subject_reference: Optional[list[SubjectReferenceItem]] = Field(
-        None,
-        description='Only available when model is S2V-01. The model will generate a video based on the subject uploaded through this parameter.',
-    )
-    duration: Optional[int] = Field(
-        None,
-        description="The length of the output video in seconds."
-    )
-    resolution: Optional[str] = Field(
-        None,
-        description="The dimensions of the video display. 1080p corresponds to 1920 x 1080 pixels, 768p corresponds to 1366 x 768 pixels."
-    )
-
-
-class MinimaxVideoGenerationResponse(BaseModel):
-    base_resp: MinimaxBaseResponse
-    task_id: str = Field(
-        ..., description='The task ID for the asynchronous video generation task.'
-    )
--- a/comfy_api_nodes/apis/pika_defs.py
+++ b/comfy_api_nodes/apis/pika_defs.py
--- a/comfy_api_nodes/apis/request_logger.py
+++ b/comfy_api_nodes/apis/request_logger.py
@@ -1,11 +1,11 @@
 from __future__ import annotations

+import os
 import datetime
-import hashlib
 import json
 import logging
-import os
 import re
+import hashlib
 from typing import Any

 import folder_paths
--- a/comfy_api_nodes/apis/topaz_api.py
+++ b/comfy_api_nodes/apis/topaz_api.py
@@ -1,133 +0,0 @@
-from typing import Optional, Union
-
-from pydantic import BaseModel, Field
-
-
-class ImageEnhanceRequest(BaseModel):
-    model: str = Field("Reimagine")
-    output_format: str = Field("jpeg")
-    subject_detection: str = Field("All")
-    face_enhancement: bool = Field(True)
-    face_enhancement_creativity: float = Field(0, description="Is ignored if face_enhancement is false")
-    face_enhancement_strength: float = Field(0.8, description="Is ignored if face_enhancement is false")
-    source_url: str = Field(...)
-    output_width: Optional[int] = Field(None)
-    output_height: Optional[int] = Field(None)
-    crop_to_fill: bool = Field(False)
-    prompt: Optional[str] = Field(None, description="Text prompt for creative upscaling guidance")
-    creativity: int = Field(3, description="Creativity settings range from 1 to 9")
-    face_preservation: str = Field("true", description="To preserve the identity of characters")
-    color_preservation: str = Field("true", description="To preserve the original color")
-
-
-class ImageAsyncTaskResponse(BaseModel):
-    process_id: str = Field(...)
-
-
-class ImageStatusResponse(BaseModel):
-    process_id: str = Field(...)
-    status: str = Field(...)
-    progress: Optional[int] = Field(None)
-    credits: int = Field(...)
-
-
-class ImageDownloadResponse(BaseModel):
-    download_url: str = Field(...)
-    expiry: int = Field(...)
-
-
-class Resolution(BaseModel):
-    width: int = Field(...)
-    height: int = Field(...)
-
-
-class CreateCreateVideoRequestSource(BaseModel):
-    container: str = Field(...)
-    size: int = Field(..., description="Size of the video file in bytes")
-    duration: int = Field(..., description="Duration of the video file in seconds")
-    frameCount: int = Field(..., description="Total number of frames in the video")
-    frameRate: int = Field(...)
-    resolution: Resolution = Field(...)
-
-
-class VideoFrameInterpolationFilter(BaseModel):
-    model: str = Field(...)
-    slowmo: Optional[int] = Field(None)
-    fps: int = Field(...)
-    duplicate: bool = Field(...)
-    duplicate_threshold: float = Field(...)
-
-
-class VideoEnhancementFilter(BaseModel):
-    model: str = Field(...)
-    auto: Optional[str] = Field(None, description="Auto, Manual, Relative")
-    focusFixLevel: Optional[str] = Field(None, description="Downscales video input for correction of blurred subjects")
-    compression: Optional[float] = Field(None, description="Strength of compression recovery")
-    details: Optional[float] = Field(None, description="Amount of detail reconstruction")
-    prenoise: Optional[float] = Field(None, description="Amount of noise to add to input to reduce over-smoothing")
-    noise: Optional[float] = Field(None, description="Amount of noise reduction")
-    halo: Optional[float] = Field(None, description="Amount of halo reduction")
-    preblur: Optional[float] = Field(None, description="Anti-aliasing and deblurring strength")
-    blur: Optional[float] = Field(None, description="Amount of sharpness applied")
-    grain: Optional[float] = Field(None, description="Grain after AI model processing")
-    grainSize: Optional[float] = Field(None, description="Size of generated grain")
-    recoverOriginalDetailValue: Optional[float] = Field(None, description="Source details into the output video")
-    creativity: Optional[str] = Field(None, description="Creativity level(high, low) for slc-1 only")
-    isOptimizedMode: Optional[bool] = Field(None, description="Set to true for Starlight Creative (slc-1) only")
-
-
-class OutputInformationVideo(BaseModel):
-    resolution: Resolution = Field(...)
-    frameRate: int = Field(...)
-    audioCodec: Optional[str] = Field(..., description="Required if audioTransfer is Copy or Convert")
-    audioTransfer: str = Field(..., description="Copy, Convert, None")
-    dynamicCompressionLevel: str = Field(..., description="Low, Mid, High")
-
-
-class Overrides(BaseModel):
-    isPaidDiffusion: bool = Field(True)
-
-
-class CreateVideoRequest(BaseModel):
-    source: CreateCreateVideoRequestSource = Field(...)
-    filters: list[Union[VideoFrameInterpolationFilter, VideoEnhancementFilter]] = Field(...)
-    output: OutputInformationVideo = Field(...)
-    overrides: Overrides = Field(Overrides(isPaidDiffusion=True))
-
-
-class CreateVideoResponse(BaseModel):
-    requestId: str = Field(...)
-
-
-class VideoAcceptResponse(BaseModel):
-    uploadId: str = Field(...)
-    urls: list[str] = Field(...)
-
-
-class VideoCompleteUploadRequestPart(BaseModel):
-    partNum: int = Field(...)
-    eTag: str = Field(...)
-
-
-class VideoCompleteUploadRequest(BaseModel):
-    uploadResults: list[VideoCompleteUploadRequestPart] = Field(...)
-
-
-class VideoCompleteUploadResponse(BaseModel):
-    message: str = Field(..., description="Confirmation message")
-
-
-class VideoStatusResponseEstimates(BaseModel):
-    cost: list[int] = Field(...)
-
-
-class VideoStatusResponseDownloadUrl(BaseModel):
-    url: str = Field(...)
-
-
-class VideoStatusResponse(BaseModel):
-    status: str = Field(...)
-    estimates: Optional[VideoStatusResponseEstimates] = Field(None)
-    progress: Optional[float] = Field(None)
-    message: Optional[str] = Field("")
-    download: Optional[VideoStatusResponseDownloadUrl] = Field(None)
--- a/comfy_api_nodes/apis/veo_api.py
+++ b/comfy_api_nodes/apis/veo_api.py
@@ -1,21 +1,34 @@
-from typing import Optional
+from typing import Optional, Union
+from enum import Enum

 from pydantic import BaseModel, Field


-class VeoRequestInstanceImage(BaseModel):
-    bytesBase64Encoded: str | None = Field(None)
-    gcsUri: str | None = Field(None)
-    mimeType: str | None = Field(None)
+class Image2(BaseModel):
+    bytesBase64Encoded: str
+    gcsUri: Optional[str] = None
+    mimeType: Optional[str] = None


-class VeoRequestInstance(BaseModel):
-    image: VeoRequestInstanceImage | None = Field(None)
-    lastFrame: VeoRequestInstanceImage | None = Field(None)
+class Image3(BaseModel):
+    bytesBase64Encoded: Optional[str] = None
+    gcsUri: str
+    mimeType: Optional[str] = None
+
+
+class Instance1(BaseModel):
+    image: Optional[Union[Image2, Image3]] = Field(
+        None, description='Optional image to guide video generation'
+    )
    prompt: str = Field(..., description='Text description of the video')


-class VeoRequestParameters(BaseModel):
+class PersonGeneration1(str, Enum):
+    ALLOW = 'ALLOW'
+    BLOCK = 'BLOCK'
+
+
+class Parameters1(BaseModel):
    aspectRatio: Optional[str] = Field(None, examples=['16:9'])
    durationSeconds: Optional[int] = None
    enhancePrompt: Optional[bool] = None
@@ -24,18 +37,17 @@ class VeoRequestParameters(BaseModel):
        description='Generate audio for the video. Only supported by veo 3 models.',
    )
    negativePrompt: Optional[str] = None
-    personGeneration: str | None = Field(None, description="ALLOW or BLOCK")
+    personGeneration: Optional[PersonGeneration1] = None
    sampleCount: Optional[int] = None
    seed: Optional[int] = None
    storageUri: Optional[str] = Field(
        None, description='Optional Cloud Storage URI to upload the video'
    )
-    resolution: str | None = Field(None)


 class VeoGenVidRequest(BaseModel):
-    instances: list[VeoRequestInstance] | None = Field(None)
-    parameters: VeoRequestParameters | None = Field(None)
+    instances: Optional[list[Instance1]] = None
+    parameters: Optional[Parameters1] = None


 class VeoGenVidResponse(BaseModel):
--- a/comfy_api_nodes/nodes_bfl.py
+++ b/comfy_api_nodes/nodes_bfl.py
@@ -1,29 +1,30 @@
 from inspect import cleandoc
+from typing import Optional

 import torch
-from pydantic import BaseModel
 from typing_extensions import override

 from comfy_api.latest import IO, ComfyExtension
+from comfy_api_nodes.apinode_utils import (
+    resize_mask_to_image,
+    validate_aspect_ratio,
+)
 from comfy_api_nodes.apis.bfl_api import (
    BFLFluxExpandImageRequest,
    BFLFluxFillImageRequest,
    BFLFluxKontextProGenerateRequest,
+    BFLFluxProGenerateRequest,
    BFLFluxProGenerateResponse,
    BFLFluxProUltraGenerateRequest,
    BFLFluxStatusResponse,
    BFLStatus,
-    Flux2ProGenerateRequest,
 )
 from comfy_api_nodes.util import (
    ApiEndpoint,
    download_url_to_image_tensor,
-    get_number_of_images,
    poll_op,
-    resize_mask_to_image,
    sync_op,
    tensor_to_base64_string,
-    validate_aspect_ratio_string,
    validate_string,
 )

@@ -42,6 +43,11 @@ class FluxProUltraImageNode(IO.ComfyNode):
    Generates images using Flux Pro 1.1 Ultra via api based on prompt and resolution.
    """

+    MINIMUM_RATIO = 1 / 4
+    MAXIMUM_RATIO = 4 / 1
+    MINIMUM_RATIO_STR = "1:4"
+    MAXIMUM_RATIO_STR = "4:1"
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
@@ -106,7 +112,16 @@ class FluxProUltraImageNode(IO.ComfyNode):

    @classmethod
    def validate_inputs(cls, aspect_ratio: str):
-        validate_aspect_ratio_string(aspect_ratio, (1, 4), (4, 1))
+        try:
+            validate_aspect_ratio(
+                aspect_ratio,
+                minimum_ratio=cls.MINIMUM_RATIO,
+                maximum_ratio=cls.MAXIMUM_RATIO,
+                minimum_ratio_str=cls.MINIMUM_RATIO_STR,
+                maximum_ratio_str=cls.MAXIMUM_RATIO_STR,
+            )
+        except Exception as e:
+            return str(e)
        return True

    @classmethod
@@ -117,7 +132,7 @@ class FluxProUltraImageNode(IO.ComfyNode):
        prompt_upsampling: bool = False,
        raw: bool = False,
        seed: int = 0,
-        image_prompt: torch.Tensor | None = None,
+        image_prompt: Optional[torch.Tensor] = None,
        image_prompt_strength: float = 0.1,
    ) -> IO.NodeOutput:
        if image_prompt is None:
@@ -130,7 +145,13 @@ class FluxProUltraImageNode(IO.ComfyNode):
                prompt=prompt,
                prompt_upsampling=prompt_upsampling,
                seed=seed,
-                aspect_ratio=aspect_ratio,
+                aspect_ratio=validate_aspect_ratio(
+                    aspect_ratio,
+                    minimum_ratio=cls.MINIMUM_RATIO,
+                    maximum_ratio=cls.MAXIMUM_RATIO,
+                    minimum_ratio_str=cls.MINIMUM_RATIO_STR,
+                    maximum_ratio_str=cls.MAXIMUM_RATIO_STR,
+                ),
                raw=raw,
                image_prompt=(image_prompt if image_prompt is None else tensor_to_base64_string(image_prompt)),
                image_prompt_strength=(None if image_prompt is None else round(image_prompt_strength, 2)),
@@ -159,6 +180,11 @@ class FluxKontextProImageNode(IO.ComfyNode):
    Edits images using Flux.1 Kontext [pro] via api based on prompt and aspect ratio.
    """

+    MINIMUM_RATIO = 1 / 4
+    MAXIMUM_RATIO = 4 / 1
+    MINIMUM_RATIO_STR = "1:4"
+    MAXIMUM_RATIO_STR = "4:1"
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
@@ -231,11 +257,17 @@ class FluxKontextProImageNode(IO.ComfyNode):
        aspect_ratio: str,
        guidance: float,
        steps: int,
-        input_image: torch.Tensor | None = None,
+        input_image: Optional[torch.Tensor] = None,
        seed=0,
        prompt_upsampling=False,
    ) -> IO.NodeOutput:
-        validate_aspect_ratio_string(aspect_ratio, (1, 4), (4, 1))
+        aspect_ratio = validate_aspect_ratio(
+            aspect_ratio,
+            minimum_ratio=cls.MINIMUM_RATIO,
+            maximum_ratio=cls.MAXIMUM_RATIO,
+            minimum_ratio_str=cls.MINIMUM_RATIO_STR,
+            maximum_ratio_str=cls.MAXIMUM_RATIO_STR,
+        )
        if input_image is None:
            validate_string(prompt, strip_whitespace=False)
        initial_response = await sync_op(
@@ -281,6 +313,124 @@ class FluxKontextMaxImageNode(FluxKontextProImageNode):
    DISPLAY_NAME = "Flux.1 Kontext [max] Image"


+class FluxProImageNode(IO.ComfyNode):
+    """
+    Generates images synchronously based on prompt and resolution.
+    """
+
+    @classmethod
+    def define_schema(cls) -> IO.Schema:
+        return IO.Schema(
+            node_id="FluxProImageNode",
+            display_name="Flux 1.1 [pro] Image",
+            category="api node/image/BFL",
+            description=cleandoc(cls.__doc__ or ""),
+            inputs=[
+                IO.String.Input(
+                    "prompt",
+                    multiline=True,
+                    default="",
+                    tooltip="Prompt for the image generation",
+                ),
+                IO.Boolean.Input(
+                    "prompt_upsampling",
+                    default=False,
+                    tooltip="Whether to perform upsampling on the prompt. "
+                    "If active, automatically modifies the prompt for more creative generation, "
+                    "but results are nondeterministic (same seed will not produce exactly the same result).",
+                ),
+                IO.Int.Input(
+                    "width",
+                    default=1024,
+                    min=256,
+                    max=1440,
+                    step=32,
+                ),
+                IO.Int.Input(
+                    "height",
+                    default=768,
+                    min=256,
+                    max=1440,
+                    step=32,
+                ),
+                IO.Int.Input(
+                    "seed",
+                    default=0,
+                    min=0,
+                    max=0xFFFFFFFFFFFFFFFF,
+                    control_after_generate=True,
+                    tooltip="The random seed used for creating the noise.",
+                ),
+                IO.Image.Input(
+                    "image_prompt",
+                    optional=True,
+                ),
+                # "image_prompt_strength": (
+                #     IO.FLOAT,
+                #     {
+                #         "default": 0.1,
+                #         "min": 0.0,
+                #         "max": 1.0,
+                #         "step": 0.01,
+                #         "tooltip": "Blend between the prompt and the image prompt.",
+                #     },
+                # ),
+            ],
+            outputs=[IO.Image.Output()],
+            hidden=[
+                IO.Hidden.auth_token_comfy_org,
+                IO.Hidden.api_key_comfy_org,
+                IO.Hidden.unique_id,
+            ],
+            is_api_node=True,
+        )
+
+    @classmethod
+    async def execute(
+        cls,
+        prompt: str,
+        prompt_upsampling,
+        width: int,
+        height: int,
+        seed=0,
+        image_prompt=None,
+        # image_prompt_strength=0.1,
+    ) -> IO.NodeOutput:
+        image_prompt = image_prompt if image_prompt is None else tensor_to_base64_string(image_prompt)
+        initial_response = await sync_op(
+            cls,
+            ApiEndpoint(
+                path="/proxy/bfl/flux-pro-1.1/generate",
+                method="POST",
+            ),
+            response_model=BFLFluxProGenerateResponse,
+            data=BFLFluxProGenerateRequest(
+                prompt=prompt,
+                prompt_upsampling=prompt_upsampling,
+                width=width,
+                height=height,
+                seed=seed,
+                image_prompt=image_prompt,
+            ),
+        )
+        response = await poll_op(
+            cls,
+            ApiEndpoint(initial_response.polling_url),
+            response_model=BFLFluxStatusResponse,
+            status_extractor=lambda r: r.status,
+            progress_extractor=lambda r: r.progress,
+            completed_statuses=[BFLStatus.ready],
+            failed_statuses=[
+                BFLStatus.request_moderated,
+                BFLStatus.content_moderated,
+                BFLStatus.error,
+                BFLStatus.task_not_found,
+            ],
+            queued_statuses=[],
+        )
+        return IO.NodeOutput(await download_url_to_image_tensor(response.result["sample"]))
+
+
 class FluxProExpandNode(IO.ComfyNode):
    """
    Outpaints image based on prompt.
@@ -523,125 +673,16 @@ class FluxProFillNode(IO.ComfyNode):
        return IO.NodeOutput(await download_url_to_image_tensor(response.result["sample"]))


-class Flux2ProImageNode(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls) -> IO.Schema:
-        return IO.Schema(
-            node_id="Flux2ProImageNode",
-            display_name="Flux.2 [pro] Image",
-            category="api node/image/BFL",
-            description="Generates images synchronously based on prompt and resolution.",
-            inputs=[
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    default="",
-                    tooltip="Prompt for the image generation or edit",
-                ),
-                IO.Int.Input(
-                    "width",
-                    default=1024,
-                    min=256,
-                    max=2048,
-                    step=32,
-                ),
-                IO.Int.Input(
-                    "height",
-                    default=768,
-                    min=256,
-                    max=2048,
-                    step=32,
-                ),
-                IO.Int.Input(
-                    "seed",
-                    default=0,
-                    min=0,
-                    max=0xFFFFFFFFFFFFFFFF,
-                    control_after_generate=True,
-                    tooltip="The random seed used for creating the noise.",
-                ),
-                IO.Boolean.Input(
-                    "prompt_upsampling",
-                    default=False,
-                    tooltip="Whether to perform upsampling on the prompt. "
-                    "If active, automatically modifies the prompt for more creative generation, "
-                    "but results are nondeterministic (same seed will not produce exactly the same result).",
-                ),
-                IO.Image.Input("images", optional=True, tooltip="Up to 4 images to be used as references."),
-            ],
-            outputs=[IO.Image.Output()],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        prompt: str,
-        width: int,
-        height: int,
-        seed: int,
-        prompt_upsampling: bool,
-        images: torch.Tensor | None = None,
-    ) -> IO.NodeOutput:
-        reference_images = {}
-        if images is not None:
-            if get_number_of_images(images) > 9:
-                raise ValueError("The current maximum number of supported images is 9.")
-            for image_index in range(images.shape[0]):
-                key_name = f"input_image_{image_index + 1}" if image_index else "input_image"
-                reference_images[key_name] = tensor_to_base64_string(images[image_index], total_pixels=2048 * 2048)
-        initial_response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/bfl/flux-2-pro/generate", method="POST"),
-            response_model=BFLFluxProGenerateResponse,
-            data=Flux2ProGenerateRequest(
-                prompt=prompt,
-                width=width,
-                height=height,
-                seed=seed,
-                prompt_upsampling=prompt_upsampling,
-                **reference_images,
-            ),
-        )
-
-        def price_extractor(_r: BaseModel) -> float | None:
-            return None if initial_response.cost is None else initial_response.cost / 100
-
-        response = await poll_op(
-            cls,
-            ApiEndpoint(initial_response.polling_url),
-            response_model=BFLFluxStatusResponse,
-            status_extractor=lambda r: r.status,
-            progress_extractor=lambda r: r.progress,
-            price_extractor=price_extractor,
-            completed_statuses=[BFLStatus.ready],
-            failed_statuses=[
-                BFLStatus.request_moderated,
-                BFLStatus.content_moderated,
-                BFLStatus.error,
-                BFLStatus.task_not_found,
-            ],
-            queued_statuses=[],
-        )
-        return IO.NodeOutput(await download_url_to_image_tensor(response.result["sample"]))
-
-
 class BFLExtension(ComfyExtension):
    @override
    async def get_node_list(self) -> list[type[IO.ComfyNode]]:
        return [
            FluxProUltraImageNode,
+            # FluxProImageNode,
            FluxKontextProImageNode,
            FluxKontextMaxImageNode,
            FluxProExpandNode,
            FluxProFillNode,
-            Flux2ProImageNode,
        ]


--- a/comfy_api_nodes/nodes_bytedance.py
+++ b/comfy_api_nodes/nodes_bytedance.py
@@ -17,7 +17,7 @@ from comfy_api_nodes.util import (
    poll_op,
    sync_op,
    upload_images_to_comfyapi,
-    validate_image_aspect_ratio,
+    validate_image_aspect_ratio_range,
    validate_image_dimensions,
    validate_string,
 )
@@ -403,7 +403,7 @@ class ByteDanceImageEditNode(IO.ComfyNode):
        validate_string(prompt, strip_whitespace=True, min_length=1)
        if get_number_of_images(image) != 1:
            raise ValueError("Exactly one input image is required.")
-        validate_image_aspect_ratio(image, (1, 3), (3, 1))
+        validate_image_aspect_ratio_range(image, (1, 3), (3, 1))
        source_url = (await upload_images_to_comfyapi(cls, image, max_images=1, mime_type="image/png"))[0]
        payload = Image2ImageTaskCreationRequest(
            model=model,
@@ -565,7 +565,7 @@ class ByteDanceSeedreamNode(IO.ComfyNode):
        reference_images_urls = []
        if n_input_images:
            for i in image:
-                validate_image_aspect_ratio(i, (1, 3), (3, 1))
+                validate_image_aspect_ratio_range(i, (1, 3), (3, 1))
            reference_images_urls = await upload_images_to_comfyapi(
                cls,
                image,
@@ -798,7 +798,7 @@ class ByteDanceImageToVideoNode(IO.ComfyNode):
        validate_string(prompt, strip_whitespace=True, min_length=1)
        raise_if_text_params(prompt, ["resolution", "ratio", "duration", "seed", "camerafixed", "watermark"])
        validate_image_dimensions(image, min_width=300, min_height=300, max_width=6000, max_height=6000)
-        validate_image_aspect_ratio(image, (2, 5), (5, 2), strict=False)  # 0.4 to 2.5
+        validate_image_aspect_ratio_range(image, (2, 5), (5, 2), strict=False)  # 0.4 to 2.5

        image_url = (await upload_images_to_comfyapi(cls, image, max_images=1))[0]
        prompt = (
@@ -923,7 +923,7 @@ class ByteDanceFirstLastFrameNode(IO.ComfyNode):
        raise_if_text_params(prompt, ["resolution", "ratio", "duration", "seed", "camerafixed", "watermark"])
        for i in (first_frame, last_frame):
            validate_image_dimensions(i, min_width=300, min_height=300, max_width=6000, max_height=6000)
-            validate_image_aspect_ratio(i, (2, 5), (5, 2), strict=False)  # 0.4 to 2.5
+            validate_image_aspect_ratio_range(i, (2, 5), (5, 2), strict=False)  # 0.4 to 2.5

        download_urls = await upload_images_to_comfyapi(
            cls,
@@ -1045,7 +1045,7 @@ class ByteDanceImageReferenceNode(IO.ComfyNode):
        raise_if_text_params(prompt, ["resolution", "ratio", "duration", "seed", "watermark"])
        for image in images:
            validate_image_dimensions(image, min_width=300, min_height=300, max_width=6000, max_height=6000)
-            validate_image_aspect_ratio(image, (2, 5), (5, 2), strict=False)  # 0.4 to 2.5
+            validate_image_aspect_ratio_range(image, (2, 5), (5, 2), strict=False)  # 0.4 to 2.5

        image_urls = await upload_images_to_comfyapi(cls, images, max_images=4, mime_type="image/png")
        prompt = (
--- a/comfy_api_nodes/nodes_gemini.py
+++ b/comfy_api_nodes/nodes_gemini.py
@@ -3,11 +3,16 @@ API Nodes for Gemini Multimodal LLM Usage via Remote API
 See: https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/inference
 """

+from __future__ import annotations
+
 import base64
+import json
 import os
+import time
+import uuid
 from enum import Enum
 from io import BytesIO
-from typing import Literal
+from typing import Literal, Optional

 import torch
 from typing_extensions import override
@@ -15,31 +20,29 @@ from typing_extensions import override
 import folder_paths
 from comfy_api.latest import IO, ComfyExtension, Input
 from comfy_api.util import VideoCodec, VideoContainer
-from comfy_api_nodes.apis.gemini_api import (
+from comfy_api_nodes.apis import (
    GeminiContent,
-    GeminiFileData,
    GeminiGenerateContentRequest,
    GeminiGenerateContentResponse,
-    GeminiImageConfig,
-    GeminiImageGenerateContentRequest,
-    GeminiImageGenerationConfig,
    GeminiInlineData,
    GeminiMimeType,
    GeminiPart,
-    GeminiRole,
-    Modality,
+)
+from comfy_api_nodes.apis.gemini_api import (
+    GeminiImageConfig,
+    GeminiImageGenerateContentRequest,
+    GeminiImageGenerationConfig,
 )
 from comfy_api_nodes.util import (
    ApiEndpoint,
    audio_to_base64_string,
    bytesio_to_image_tensor,
-    get_number_of_images,
    sync_op,
    tensor_to_base64_string,
-    upload_images_to_comfyapi,
    validate_string,
    video_to_base64_string,
 )
+from server import PromptServer

 GEMINI_BASE_ENDPOINT = "/proxy/vertexai/gemini"
 GEMINI_MAX_INPUT_FILE_SIZE = 20 * 1024 * 1024  # 20 MB
@@ -54,7 +57,6 @@ class GeminiModel(str, Enum):
    gemini_2_5_flash_preview_04_17 = "gemini-2.5-flash-preview-04-17"
    gemini_2_5_pro = "gemini-2.5-pro"
    gemini_2_5_flash = "gemini-2.5-flash"
-    gemini_3_0_pro = "gemini-3-pro-preview"


 class GeminiImageModel(str, Enum):
@@ -66,43 +68,24 @@ class GeminiImageModel(str, Enum):
    gemini_2_5_flash_image = "gemini-2.5-flash-image"


-async def create_image_parts(
-    cls: type[IO.ComfyNode],
-    images: torch.Tensor,
-    image_limit: int = 0,
-) -> list[GeminiPart]:
+def create_image_parts(image_input: torch.Tensor) -> list[GeminiPart]:
+    """
+    Convert image tensor input to Gemini API compatible parts.
+
+    Args:
+        image_input: Batch of image tensors from ComfyUI.
+
+    Returns:
+        List of GeminiPart objects containing the encoded images.
+    """
    image_parts: list[GeminiPart] = []
-    if image_limit < 0:
-        raise ValueError("image_limit must be greater than or equal to 0 when creating Gemini image parts.")
-    total_images = get_number_of_images(images)
-    if total_images <= 0:
-        raise ValueError("No images provided to create_image_parts; at least one image is required.")
-
-    # If image_limit == 0 --> use all images; otherwise clamp to image_limit.
-    effective_max = total_images if image_limit == 0 else min(total_images, image_limit)
-
-    # Number of images we'll send as URLs (fileData)
-    num_url_images = min(effective_max, 10)  # Vertex API max number of image links
-    reference_images_urls = await upload_images_to_comfyapi(
-        cls,
-        images,
-        max_images=num_url_images,
-    )
-    for reference_image_url in reference_images_urls:
-        image_parts.append(
-            GeminiPart(
-                fileData=GeminiFileData(
-                    mimeType=GeminiMimeType.image_png,
-                    fileUri=reference_image_url,
-                )
-            )
-        )
-    for idx in range(num_url_images, effective_max):
+    for image_index in range(image_input.shape[0]):
+        image_as_b64 = tensor_to_base64_string(image_input[image_index].unsqueeze(0))
        image_parts.append(
            GeminiPart(
                inlineData=GeminiInlineData(
                    mimeType=GeminiMimeType.image_png,
-                    data=tensor_to_base64_string(images[idx]),
+                    data=image_as_b64,
                )
            )
        )
@@ -120,16 +103,6 @@ def get_parts_by_type(response: GeminiGenerateContentResponse, part_type: Litera
    Returns:
        List of response parts matching the requested type.
    """
-    if response.candidates is None:
-        if response.promptFeedback and response.promptFeedback.blockReason:
-            feedback = response.promptFeedback
-            raise ValueError(
-                f"Gemini API blocked the request. Reason: {feedback.blockReason} ({feedback.blockReasonMessage})"
-            )
-        raise ValueError(
-            "Gemini API returned no response candidates. If you are using the `IMAGE` modality, "
-            "try changing it to `IMAGE+TEXT` to view the model's reasoning and understand why image generation failed."
-        )
    parts = []
    for part in response.candidates[0].content.parts:
        if part_type == "text" and hasattr(part, "text") and part.text:
@@ -166,50 +139,6 @@ def get_image_from_response(response: GeminiGenerateContentResponse) -> torch.Te
    return torch.cat(image_tensors, dim=0)


-def calculate_tokens_price(response: GeminiGenerateContentResponse) -> float | None:
-    if not response.modelVersion:
-        return None
-    # Define prices (Cost per 1,000,000 tokens), see https://cloud.google.com/vertex-ai/generative-ai/pricing
-    if response.modelVersion in ("gemini-2.5-pro-preview-05-06", "gemini-2.5-pro"):
-        input_tokens_price = 1.25
-        output_text_tokens_price = 10.0
-        output_image_tokens_price = 0.0
-    elif response.modelVersion in (
-        "gemini-2.5-flash-preview-04-17",
-        "gemini-2.5-flash",
-    ):
-        input_tokens_price = 0.30
-        output_text_tokens_price = 2.50
-        output_image_tokens_price = 0.0
-    elif response.modelVersion in (
-        "gemini-2.5-flash-image-preview",
-        "gemini-2.5-flash-image",
-    ):
-        input_tokens_price = 0.30
-        output_text_tokens_price = 2.50
-        output_image_tokens_price = 30.0
-    elif response.modelVersion == "gemini-3-pro-preview":
-        input_tokens_price = 2
-        output_text_tokens_price = 12.0
-        output_image_tokens_price = 0.0
-    elif response.modelVersion == "gemini-3-pro-image-preview":
-        input_tokens_price = 2
-        output_text_tokens_price = 12.0
-        output_image_tokens_price = 120.0
-    else:
-        return None
-    final_price = response.usageMetadata.promptTokenCount * input_tokens_price
-    if response.usageMetadata.candidatesTokensDetails:
-        for i in response.usageMetadata.candidatesTokensDetails:
-            if i.modality == Modality.IMAGE:
-                final_price += output_image_tokens_price * i.tokenCount  # for Nano Banana models
-            else:
-                final_price += output_text_tokens_price * i.tokenCount
-    if response.usageMetadata.thoughtsTokenCount:
-        final_price += output_text_tokens_price * response.usageMetadata.thoughtsTokenCount
-    return final_price / 1_000_000.0
-
-
 class GeminiNode(IO.ComfyNode):
    """
    Node to generate text responses from a Gemini model.
@@ -343,10 +272,10 @@ class GeminiNode(IO.ComfyNode):
        prompt: str,
        model: str,
        seed: int,
-        images: torch.Tensor | None = None,
-        audio: Input.Audio | None = None,
-        video: Input.Video | None = None,
-        files: list[GeminiPart] | None = None,
+        images: Optional[torch.Tensor] = None,
+        audio: Optional[Input.Audio] = None,
+        video: Optional[Input.Video] = None,
+        files: Optional[list[GeminiPart]] = None,
    ) -> IO.NodeOutput:
        validate_string(prompt, strip_whitespace=False)

@@ -355,7 +284,8 @@ class GeminiNode(IO.ComfyNode):

        # Add other modal parts
        if images is not None:
-            parts.extend(await create_image_parts(cls, images))
+            image_parts = create_image_parts(images)
+            parts.extend(image_parts)
        if audio is not None:
            parts.extend(cls.create_audio_parts(audio))
        if video is not None:
@@ -370,16 +300,39 @@ class GeminiNode(IO.ComfyNode):
            data=GeminiGenerateContentRequest(
                contents=[
                    GeminiContent(
-                        role=GeminiRole.user,
+                        role="user",
                        parts=parts,
                    )
                ]
            ),
            response_model=GeminiGenerateContentResponse,
-            price_extractor=calculate_tokens_price,
        )

+        # Get result output
        output_text = get_text_from_response(response)
+        if output_text:
+            # Not a true chat history like the OpenAI Chat node. It is emulated so the frontend can show a copy button.
+            render_spec = {
+                "node_id": cls.hidden.unique_id,
+                "component": "ChatHistoryWidget",
+                "props": {
+                    "history": json.dumps(
+                        [
+                            {
+                                "prompt": prompt,
+                                "response": output_text,
+                                "response_id": str(uuid.uuid4()),
+                                "timestamp": time.time(),
+                            }
+                        ]
+                    ),
+                },
+            }
+            PromptServer.instance.send_sync(
+                "display_component",
+                render_spec,
+            )
+
        return IO.NodeOutput(output_text or "Empty response from Gemini model...")


@@ -453,7 +406,7 @@ class GeminiInputFiles(IO.ComfyNode):
        )

    @classmethod
-    def execute(cls, file: str, GEMINI_INPUT_FILES: list[GeminiPart] | None = None) -> IO.NodeOutput:
+    def execute(cls, file: str, GEMINI_INPUT_FILES: Optional[list[GeminiPart]] = None) -> IO.NodeOutput:
        """Loads and formats input files for Gemini API."""
        if GEMINI_INPUT_FILES is None:
            GEMINI_INPUT_FILES = []
@@ -468,7 +421,7 @@ class GeminiImage(IO.ComfyNode):
    def define_schema(cls):
        return IO.Schema(
            node_id="GeminiImageNode",
-            display_name="Nano Banana (Google Gemini Image)",
+            display_name="Google Gemini Image",
            category="api node/image/Gemini",
            description="Edit images synchronously via Google API.",
            inputs=[
@@ -516,13 +469,6 @@ class GeminiImage(IO.ComfyNode):
                    "or otherwise generates 1:1 squares.",
                    optional=True,
                ),
-                IO.Combo.Input(
-                    "response_modalities",
-                    options=["IMAGE+TEXT", "IMAGE"],
-                    tooltip="Choose 'IMAGE' for image-only output, or "
-                    "'IMAGE+TEXT' to return both the generated image and a text response.",
-                    optional=True,
-                ),
            ],
            outputs=[
                IO.Image.Output(),
@@ -542,10 +488,9 @@ class GeminiImage(IO.ComfyNode):
        prompt: str,
        model: str,
        seed: int,
-        images: torch.Tensor | None = None,
-        files: list[GeminiPart] | None = None,
+        images: Optional[torch.Tensor] = None,
+        files: Optional[list[GeminiPart]] = None,
        aspect_ratio: str = "auto",
-        response_modalities: str = "IMAGE+TEXT",
    ) -> IO.NodeOutput:
        validate_string(prompt, strip_whitespace=True, min_length=1)
        parts: list[GeminiPart] = [GeminiPart(text=prompt)]
@@ -555,7 +500,8 @@ class GeminiImage(IO.ComfyNode):
        image_config = GeminiImageConfig(aspectRatio=aspect_ratio)

        if images is not None:
-            parts.extend(await create_image_parts(cls, images))
+            image_parts = create_image_parts(images)
+            parts.extend(image_parts)
        if files is not None:
            parts.extend(files)

@@ -564,137 +510,43 @@ class GeminiImage(IO.ComfyNode):
            endpoint=ApiEndpoint(path=f"{GEMINI_BASE_ENDPOINT}/{model}", method="POST"),
            data=GeminiImageGenerateContentRequest(
                contents=[
-                    GeminiContent(role=GeminiRole.user, parts=parts),
+                    GeminiContent(role="user", parts=parts),
                ],
                generationConfig=GeminiImageGenerationConfig(
-                    responseModalities=(["IMAGE"] if response_modalities == "IMAGE" else ["TEXT", "IMAGE"]),
+                    responseModalities=["TEXT", "IMAGE"],
                    imageConfig=None if aspect_ratio == "auto" else image_config,
                ),
            ),
            response_model=GeminiGenerateContentResponse,
-            price_extractor=calculate_tokens_price,
-        )
-        return IO.NodeOutput(get_image_from_response(response), get_text_from_response(response))
-
-
-class GeminiImage2(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls):
-        return IO.Schema(
-            node_id="GeminiImage2Node",
-            display_name="Nano Banana Pro (Google Gemini Image)",
-            category="api node/image/Gemini",
-            description="Generate or edit images synchronously via Google Vertex API.",
-            inputs=[
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    tooltip="Text prompt describing the image to generate or the edits to apply. "
-                    "Include any constraints, styles, or details the model should follow.",
-                    default="",
-                ),
-                IO.Combo.Input(
-                    "model",
-                    options=["gemini-3-pro-image-preview"],
-                ),
-                IO.Int.Input(
-                    "seed",
-                    default=42,
-                    min=0,
-                    max=0xFFFFFFFFFFFFFFFF,
-                    control_after_generate=True,
-                    tooltip="When the seed is fixed to a specific value, the model makes a best effort to provide "
-                    "the same response for repeated requests. Deterministic output isn't guaranteed. "
-                    "Also, changing the model or parameter settings, such as the temperature, "
-                    "can cause variations in the response even when you use the same seed value. "
-                    "By default, a random seed value is used.",
-                ),
-                IO.Combo.Input(
-                    "aspect_ratio",
-                    options=["auto", "1:1", "2:3", "3:2", "3:4", "4:3", "4:5", "5:4", "9:16", "16:9", "21:9"],
-                    default="auto",
-                    tooltip="If set to 'auto', matches your input image's aspect ratio; "
-                    "if no image is provided, a 16:9 square is usually generated.",
-                ),
-                IO.Combo.Input(
-                    "resolution",
-                    options=["1K", "2K", "4K"],
-                    tooltip="Target output resolution. For 2K/4K the native Gemini upscaler is used.",
-                ),
-                IO.Combo.Input(
-                    "response_modalities",
-                    options=["IMAGE+TEXT", "IMAGE"],
-                    tooltip="Choose 'IMAGE' for image-only output, or "
-                    "'IMAGE+TEXT' to return both the generated image and a text response.",
-                ),
-                IO.Image.Input(
-                    "images",
-                    optional=True,
-                    tooltip="Optional reference image(s). "
-                    "To include multiple images, use the Batch Images node (up to 14).",
-                ),
-                IO.Custom("GEMINI_INPUT_FILES").Input(
-                    "files",
-                    optional=True,
-                    tooltip="Optional file(s) to use as context for the model. "
-                    "Accepts inputs from the Gemini Generate Content Input Files node.",
-                ),
-            ],
-            outputs=[
-                IO.Image.Output(),
-                IO.String.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
        )

-    @classmethod
-    async def execute(
-        cls,
-        prompt: str,
-        model: str,
-        seed: int,
-        aspect_ratio: str,
-        resolution: str,
-        response_modalities: str,
-        images: torch.Tensor | None = None,
-        files: list[GeminiPart] | None = None,
-    ) -> IO.NodeOutput:
-        validate_string(prompt, strip_whitespace=True, min_length=1)
+        output_image = get_image_from_response(response)
+        output_text = get_text_from_response(response)
+        if output_text:
+            # Not a true chat history like the OpenAI Chat node. It is emulated so the frontend can show a copy button.
+            render_spec = {
+                "node_id": cls.hidden.unique_id,
+                "component": "ChatHistoryWidget",
+                "props": {
+                    "history": json.dumps(
+                        [
+                            {
+                                "prompt": prompt,
+                                "response": output_text,
+                                "response_id": str(uuid.uuid4()),
+                                "timestamp": time.time(),
+                            }
+                        ]
+                    ),
+                },
+            }
+            PromptServer.instance.send_sync(
+                "display_component",
+                render_spec,
+            )

-        parts: list[GeminiPart] = [GeminiPart(text=prompt)]
-        if images is not None:
-            if get_number_of_images(images) > 14:
-                raise ValueError("The current maximum number of supported images is 14.")
-            parts.extend(await create_image_parts(cls, images))
-        if files is not None:
-            parts.extend(files)
-
-        image_config = GeminiImageConfig(imageSize=resolution)
-        if aspect_ratio != "auto":
-            image_config.aspectRatio = aspect_ratio
-
-        response = await sync_op(
-            cls,
-            ApiEndpoint(path=f"{GEMINI_BASE_ENDPOINT}/{model}", method="POST"),
-            data=GeminiImageGenerateContentRequest(
-                contents=[
-                    GeminiContent(role=GeminiRole.user, parts=parts),
-                ],
-                generationConfig=GeminiImageGenerationConfig(
-                    responseModalities=(["IMAGE"] if response_modalities == "IMAGE" else ["TEXT", "IMAGE"]),
-                    imageConfig=image_config,
-                ),
-            ),
-            response_model=GeminiGenerateContentResponse,
-            price_extractor=calculate_tokens_price,
-        )
-        return IO.NodeOutput(get_image_from_response(response), get_text_from_response(response))
+        output_text = output_text or "Empty response from Gemini model..."
+        return IO.NodeOutput(output_image, output_text)


 class GeminiExtension(ComfyExtension):
@@ -703,7 +555,6 @@ class GeminiExtension(ComfyExtension):
        return [
            GeminiNode,
            GeminiImage,
-            GeminiImage2,
            GeminiInputFiles,
        ]

--- a/comfy_api_nodes/nodes_ideogram.py
+++ b/comfy_api_nodes/nodes_ideogram.py
@@ -1,6 +1,6 @@
 from io import BytesIO
 from typing_extensions import override
-from comfy_api.latest import IO, ComfyExtension
+from comfy_api.latest import ComfyExtension, IO
 from PIL import Image
 import numpy as np
 import torch
@@ -11,14 +11,20 @@ from comfy_api_nodes.apis import (
    IdeogramV3Request,
    IdeogramV3EditRequest,
 )
-from comfy_api_nodes.util import (
+
+from comfy_api_nodes.apis.client import (
    ApiEndpoint,
-    bytesio_to_image_tensor,
-    download_url_as_bytesio,
-    resize_mask_to_image,
-    sync_op,
+    HttpMethod,
+    SynchronousOperation,
 )

+from comfy_api_nodes.apinode_utils import (
+    download_url_to_bytesio,
+    bytesio_to_image_tensor,
+    resize_mask_to_image,
+)
+from server import PromptServer
+
 V1_V1_RES_MAP = {
  "Auto":"AUTO",
  "512 x 1536":"RESOLUTION_512_1536",
@@ -214,7 +220,7 @@ async def download_and_process_images(image_urls):

    for image_url in image_urls:
        # Using functions from apinode_utils.py to handle downloading and processing
-        image_bytesio = await download_url_as_bytesio(image_url)  # Download image content to BytesIO
+        image_bytesio = await download_url_to_bytesio(image_url)  # Download image content to BytesIO
        img_tensor = bytesio_to_image_tensor(image_bytesio, mode="RGB")  # Convert to torch.Tensor with RGB mode
        image_tensors.append(img_tensor)

@@ -227,6 +233,19 @@ async def download_and_process_images(image_urls):
    return stacked_tensors


+def display_image_urls_on_node(image_urls, node_id):
+    if node_id and image_urls:
+        if len(image_urls) == 1:
+            PromptServer.instance.send_progress_text(
+                f"Generated Image URL:\n{image_urls[0]}", node_id
+            )
+        else:
+            urls_text = "Generated Image URLs:\n" + "\n".join(
+                f"{i+1}. {url}" for i, url in enumerate(image_urls)
+            )
+            PromptServer.instance.send_progress_text(urls_text, node_id)
+
+
 class IdeogramV1(IO.ComfyNode):

    @classmethod
@@ -315,30 +334,44 @@ class IdeogramV1(IO.ComfyNode):
        aspect_ratio = V1_V2_RATIO_MAP.get(aspect_ratio, None)
        model = "V_1_TURBO" if turbo else "V_1"

-        response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/ideogram/generate", method="POST"),
-            response_model=IdeogramGenerateResponse,
-            data=IdeogramGenerateRequest(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/ideogram/generate",
+                method=HttpMethod.POST,
+                request_model=IdeogramGenerateRequest,
+                response_model=IdeogramGenerateResponse,
+            ),
+            request=IdeogramGenerateRequest(
                image_request=ImageRequest(
                    prompt=prompt,
                    model=model,
                    num_images=num_images,
                    seed=seed,
                    aspect_ratio=aspect_ratio if aspect_ratio != "ASPECT_1_1" else None,
-                    magic_prompt_option=(magic_prompt_option if magic_prompt_option != "AUTO" else None),
+                    magic_prompt_option=(
+                        magic_prompt_option if magic_prompt_option != "AUTO" else None
+                    ),
                    negative_prompt=negative_prompt if negative_prompt else None,
                )
            ),
-            max_retries=1,
+            auth_kwargs=auth,
        )

+        response = await operation.execute()
+
        if not response.data or len(response.data) == 0:
            raise Exception("No images were generated in the response")

        image_urls = [image_data.url for image_data in response.data if image_data.url]
+
        if not image_urls:
            raise Exception("No image URLs were generated in the response")
+
+        display_image_urls_on_node(image_urls, cls.hidden.unique_id)
        return IO.NodeOutput(await download_and_process_images(image_urls))


@@ -467,11 +500,18 @@ class IdeogramV2(IO.ComfyNode):
        else:
            final_aspect_ratio = aspect_ratio if aspect_ratio != "ASPECT_1_1" else None

-        response = await sync_op(
-            cls,
-            endpoint=ApiEndpoint(path="/proxy/ideogram/generate", method="POST"),
-            response_model=IdeogramGenerateResponse,
-            data=IdeogramGenerateRequest(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/ideogram/generate",
+                method=HttpMethod.POST,
+                request_model=IdeogramGenerateRequest,
+                response_model=IdeogramGenerateResponse,
+            ),
+            request=IdeogramGenerateRequest(
                image_request=ImageRequest(
                    prompt=prompt,
                    model=model,
@@ -479,20 +519,28 @@ class IdeogramV2(IO.ComfyNode):
                    seed=seed,
                    aspect_ratio=final_aspect_ratio,
                    resolution=final_resolution,
-                    magic_prompt_option=(magic_prompt_option if magic_prompt_option != "AUTO" else None),
+                    magic_prompt_option=(
+                        magic_prompt_option if magic_prompt_option != "AUTO" else None
+                    ),
                    style_type=style_type if style_type != "NONE" else None,
                    negative_prompt=negative_prompt if negative_prompt else None,
                    color_palette=color_palette if color_palette else None,
                )
            ),
-            max_retries=1,
+            auth_kwargs=auth,
        )
+
+        response = await operation.execute()
+
        if not response.data or len(response.data) == 0:
            raise Exception("No images were generated in the response")

        image_urls = [image_data.url for image_data in response.data if image_data.url]
+
        if not image_urls:
            raise Exception("No image URLs were generated in the response")
+
+        display_image_urls_on_node(image_urls, cls.hidden.unique_id)
        return IO.NodeOutput(await download_and_process_images(image_urls))


@@ -608,6 +656,10 @@ class IdeogramV3(IO.ComfyNode):
        character_image=None,
        character_mask=None,
    ):
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        if rendering_speed == "BALANCED":  # for backward compatibility
            rendering_speed = "DEFAULT"

@@ -642,6 +694,9 @@ class IdeogramV3(IO.ComfyNode):

        # Check if both image and mask are provided for editing mode
        if image is not None and mask is not None:
+            # Edit mode
+            path = "/proxy/ideogram/ideogram-v3/edit"
+
            # Process image and mask
            input_tensor = image.squeeze().cpu()
            # Resize mask to match image dimension
@@ -694,20 +749,27 @@ class IdeogramV3(IO.ComfyNode):
            if character_mask_binary:
                files["character_mask_binary"] = character_mask_binary

-            response = await sync_op(
-                cls,
-                ApiEndpoint(path="/proxy/ideogram/ideogram-v3/edit", method="POST"),
-                response_model=IdeogramGenerateResponse,
-                data=edit_request,
+            # Execute the operation for edit mode
+            operation = SynchronousOperation(
+                endpoint=ApiEndpoint(
+                    path=path,
+                    method=HttpMethod.POST,
+                    request_model=IdeogramV3EditRequest,
+                    response_model=IdeogramGenerateResponse,
+                ),
+                request=edit_request,
                files=files,
                content_type="multipart/form-data",
-                max_retries=1,
+                auth_kwargs=auth,
            )

        elif image is not None or mask is not None:
            # If only one of image or mask is provided, raise an error
            raise Exception("Ideogram V3 image editing requires both an image AND a mask")
        else:
+            # Generation mode
+            path = "/proxy/ideogram/ideogram-v3/generate"
+
            # Create generation request
            gen_request = IdeogramV3Request(
                prompt=prompt,
@@ -738,22 +800,32 @@ class IdeogramV3(IO.ComfyNode):
            if files:
                gen_request.style_type = "AUTO"

-            response = await sync_op(
-                cls,
-                endpoint=ApiEndpoint(path="/proxy/ideogram/ideogram-v3/generate", method="POST"),
-                response_model=IdeogramGenerateResponse,
-                data=gen_request,
+            # Execute the operation for generation mode
+            operation = SynchronousOperation(
+                endpoint=ApiEndpoint(
+                    path=path,
+                    method=HttpMethod.POST,
+                    request_model=IdeogramV3Request,
+                    response_model=IdeogramGenerateResponse,
+                ),
+                request=gen_request,
                files=files if files else None,
                content_type="multipart/form-data",
-                max_retries=1,
+                auth_kwargs=auth,
            )

+        # Execute the operation and process response
+        response = await operation.execute()
+
        if not response.data or len(response.data) == 0:
            raise Exception("No images were generated in the response")

        image_urls = [image_data.url for image_data in response.data if image_data.url]
+
        if not image_urls:
            raise Exception("No image URLs were generated in the response")
+
+        display_image_urls_on_node(image_urls, cls.hidden.unique_id)
        return IO.NodeOutput(await download_and_process_images(image_urls))


@@ -766,6 +838,5 @@ class IdeogramExtension(ComfyExtension):
            IdeogramV3,
        ]

-
 async def comfy_entrypoint() -> IdeogramExtension:
    return IdeogramExtension()
--- a/comfy_api_nodes/nodes_kling.py
+++ b/comfy_api_nodes/nodes_kling.py
@@ -4,13 +4,15 @@ For source of truth on the allowed permutations of request fields, please refere
 - [Compatibility Table](https://app.klingai.com/global/dev/document-api/apiReference/model/skillsMap)
 """

-import logging
+from __future__ import annotations
+from typing import Optional, TypeVar
 import math
+import logging

-import torch
 from typing_extensions import override

-from comfy_api.latest import IO, ComfyExtension, Input, InputImpl
+import torch
+
 from comfy_api_nodes.apis import (
    KlingCameraControl,
    KlingCameraConfig,
@@ -48,31 +50,25 @@ from comfy_api_nodes.apis import (
    KlingCharacterEffectModelName,
    KlingSingleImageEffectModelName,
 )
-from comfy_api_nodes.apis.kling_api import (
-    OmniParamImage,
-    OmniParamVideo,
-    OmniProFirstLastFrameRequest,
-    OmniProReferences2VideoRequest,
-    OmniProText2VideoRequest,
-    TaskStatusVideoResponse,
-)
 from comfy_api_nodes.util import (
-    ApiEndpoint,
-    download_url_to_image_tensor,
-    download_url_to_video_output,
-    get_number_of_images,
-    poll_op,
-    sync_op,
-    tensor_to_base64_string,
-    upload_audio_to_comfyapi,
-    upload_images_to_comfyapi,
-    upload_video_to_comfyapi,
-    validate_image_aspect_ratio,
    validate_image_dimensions,
-    validate_string,
+    validate_image_aspect_ratio,
    validate_video_dimensions,
    validate_video_duration,
+    tensor_to_base64_string,
+    validate_string,
+    upload_audio_to_comfyapi,
+    download_url_to_image_tensor,
+    upload_video_to_comfyapi,
+    download_url_to_video_output,
+    sync_op,
+    ApiEndpoint,
+    poll_op,
 )
+from comfy_api.input_impl import VideoFromFile
+from comfy_api.input.basic_types import AudioInput
+from comfy_api.input.video_types import VideoInput
+from comfy_api.latest import ComfyExtension, IO

 KLING_API_VERSION = "v1"
 PATH_TEXT_TO_VIDEO = f"/proxy/kling/{KLING_API_VERSION}/videos/text2video"
@@ -98,6 +94,8 @@ AVERAGE_DURATION_IMAGE_GEN = 32
 AVERAGE_DURATION_VIDEO_EFFECTS = 320
 AVERAGE_DURATION_VIDEO_EXTEND = 320

+R = TypeVar("R")
+

 MODE_TEXT2VIDEO = {
    "standard mode / 5s duration / kling-v1": ("std", "5", "kling-v1"),
@@ -132,8 +130,6 @@ MODE_START_END_FRAME = {
    "pro mode / 10s duration / kling-v1-6": ("pro", "10", "kling-v1-6"),
    "pro mode / 5s duration / kling-v2-1": ("pro", "5", "kling-v2-1"),
    "pro mode / 10s duration / kling-v2-1": ("pro", "10", "kling-v2-1"),
-    "pro mode / 5s duration / kling-v2-5-turbo": ("pro", "5", "kling-v2-5-turbo"),
-    "pro mode / 10s duration / kling-v2-5-turbo": ("pro", "10", "kling-v2-5-turbo"),
 }
 """
 Returns a mapping of mode strings to their corresponding (mode, duration, model_name) tuples.
@@ -210,20 +206,6 @@ VOICES_CONFIG = {
 }


-async def finish_omni_video_task(cls: type[IO.ComfyNode], response: TaskStatusVideoResponse) -> IO.NodeOutput:
-    if response.code:
-        raise RuntimeError(
-            f"Kling request failed. Code: {response.code}, Message: {response.message}, Data: {response.data}"
-        )
-    final_response = await poll_op(
-        cls,
-        ApiEndpoint(path=f"/proxy/kling/v1/videos/omni-video/{response.data.task_id}"),
-        response_model=TaskStatusVideoResponse,
-        status_extractor=lambda r: (r.data.task_status if r.data else None),
-    )
-    return IO.NodeOutput(await download_url_to_video_output(final_response.data.task_result.videos[0].url))
-
-
 def is_valid_camera_control_configs(configs: list[float]) -> bool:
    """Verifies that at least one camera control configuration is non-zero."""
    return any(not math.isclose(value, 0.0) for value in configs)
@@ -300,7 +282,7 @@ def validate_input_image(image: torch.Tensor) -> None:
    See: https://app.klingai.com/global/dev/document-api/apiReference/model/imageToVideo
    """
    validate_image_dimensions(image, min_width=300, min_height=300)
-    validate_image_aspect_ratio(image, (1, 2.5), (2.5, 1))
+    validate_image_aspect_ratio(image, min_aspect_ratio=1 / 2.5, max_aspect_ratio=2.5)


 def get_video_from_response(response) -> KlingVideoResult:
@@ -314,7 +296,7 @@ def get_video_from_response(response) -> KlingVideoResult:
    return video


-def get_video_url_from_response(response) -> str | None:
+def get_video_url_from_response(response) -> Optional[str]:
    """Returns the first video url from the Kling video generation task result.
    Will not raise an error if the response is not valid.
    """
@@ -333,7 +315,7 @@ def get_images_from_response(response) -> list[KlingImageResult]:
    return images


-def get_images_urls_from_response(response) -> str | None:
+def get_images_urls_from_response(response) -> Optional[str]:
    """Returns the list of image urls from the Kling image generation task result.
    Will not raise an error if the response is not valid. If there is only one image, returns the url as a string. If there are multiple images, returns a list of urls.
    """
@@ -367,7 +349,7 @@ async def execute_text2video(
    model_mode: str,
    duration: str,
    aspect_ratio: str,
-    camera_control: KlingCameraControl | None = None,
+    camera_control: Optional[KlingCameraControl] = None,
 ) -> IO.NodeOutput:
    validate_prompts(prompt, negative_prompt, MAX_PROMPT_LENGTH_T2V)
    task_creation_response = await sync_op(
@@ -412,8 +394,8 @@ async def execute_image2video(
    model_mode: str,
    aspect_ratio: str,
    duration: str,
-    camera_control: KlingCameraControl | None = None,
-    end_frame: torch.Tensor | None = None,
+    camera_control: Optional[KlingCameraControl] = None,
+    end_frame: Optional[torch.Tensor] = None,
 ) -> IO.NodeOutput:
    validate_prompts(prompt, negative_prompt, MAX_PROMPT_LENGTH_I2V)
    validate_input_image(start_frame)
@@ -469,9 +451,9 @@ async def execute_video_effect(
    model_name: str,
    duration: KlingVideoGenDuration,
    image_1: torch.Tensor,
-    image_2: torch.Tensor | None = None,
-    model_mode: KlingVideoGenMode | None = None,
-) -> tuple[InputImpl.VideoFromFile, str, str]:
+    image_2: Optional[torch.Tensor] = None,
+    model_mode: Optional[KlingVideoGenMode] = None,
+) -> tuple[VideoFromFile, str, str]:
    if dual_character:
        request_input_field = KlingDualCharacterEffectInput(
            model_name=model_name,
@@ -517,13 +499,13 @@ async def execute_video_effect(

 async def execute_lipsync(
    cls: type[IO.ComfyNode],
-    video: Input.Video,
-    audio: Input.Audio | None = None,
-    voice_language: str | None = None,
-    model_mode: str | None = None,
-    text: str | None = None,
-    voice_speed: float | None = None,
-    voice_id: str | None = None,
+    video: VideoInput,
+    audio: Optional[AudioInput] = None,
+    voice_language: Optional[str] = None,
+    model_mode: Optional[str] = None,
+    text: Optional[str] = None,
+    voice_speed: Optional[float] = None,
+    voice_id: Optional[str] = None,
 ) -> IO.NodeOutput:
    if text:
        validate_string(text, field_name="Text", max_length=MAX_PROMPT_LENGTH_LIP_SYNC)
@@ -536,9 +518,7 @@ async def execute_lipsync(

    # Upload the audio file to Comfy API and get download URL
    if audio:
-        audio_url = await upload_audio_to_comfyapi(
-            cls, audio, container_format="mp3", codec_name="libmp3lame", mime_type="audio/mpeg", filename="output.mp3"
-        )
+        audio_url = await upload_audio_to_comfyapi(cls, audio)
        logging.info("Uploaded audio to Comfy API. URL: %s", audio_url)
    else:
        audio_url = None
@@ -758,386 +738,6 @@ class KlingTextToVideoNode(IO.ComfyNode):
        )


-class OmniProTextToVideoNode(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls) -> IO.Schema:
-        return IO.Schema(
-            node_id="KlingOmniProTextToVideoNode",
-            display_name="Kling Omni Text to Video (Pro)",
-            category="api node/video/Kling",
-            description="Use text prompts to generate videos with the latest Kling model.",
-            inputs=[
-                IO.Combo.Input("model_name", options=["kling-video-o1"]),
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    tooltip="A text prompt describing the video content. "
-                    "This can include both positive and negative descriptions.",
-                ),
-                IO.Combo.Input("aspect_ratio", options=["16:9", "9:16", "1:1"]),
-                IO.Combo.Input("duration", options=[5, 10]),
-            ],
-            outputs=[
-                IO.Video.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        model_name: str,
-        prompt: str,
-        aspect_ratio: str,
-        duration: int,
-    ) -> IO.NodeOutput:
-        validate_string(prompt, min_length=1, max_length=2500)
-        response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/kling/v1/videos/omni-video", method="POST"),
-            response_model=TaskStatusVideoResponse,
-            data=OmniProText2VideoRequest(
-                model_name=model_name,
-                prompt=prompt,
-                aspect_ratio=aspect_ratio,
-                duration=str(duration),
-            ),
-        )
-        return await finish_omni_video_task(cls, response)
-
-
-class OmniProFirstLastFrameNode(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls) -> IO.Schema:
-        return IO.Schema(
-            node_id="KlingOmniProFirstLastFrameNode",
-            display_name="Kling Omni First-Last-Frame to Video (Pro)",
-            category="api node/video/Kling",
-            description="Use a start frame, an optional end frame, or reference images with the latest Kling model.",
-            inputs=[
-                IO.Combo.Input("model_name", options=["kling-video-o1"]),
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    tooltip="A text prompt describing the video content. "
-                    "This can include both positive and negative descriptions.",
-                ),
-                IO.Combo.Input("duration", options=["5", "10"]),
-                IO.Image.Input("first_frame"),
-                IO.Image.Input(
-                    "end_frame",
-                    optional=True,
-                    tooltip="An optional end frame for the video. "
-                    "This cannot be used simultaneously with 'reference_images'.",
-                ),
-                IO.Image.Input(
-                    "reference_images",
-                    optional=True,
-                    tooltip="Up to 6 additional reference images.",
-                ),
-            ],
-            outputs=[
-                IO.Video.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        model_name: str,
-        prompt: str,
-        duration: int,
-        first_frame: Input.Image,
-        end_frame: Input.Image | None = None,
-        reference_images: Input.Image | None = None,
-    ) -> IO.NodeOutput:
-        validate_string(prompt, min_length=1, max_length=2500)
-        if end_frame is not None and reference_images is not None:
-            raise ValueError("The 'end_frame' input cannot be used simultaneously with 'reference_images'.")
-        validate_image_dimensions(first_frame, min_width=300, min_height=300)
-        validate_image_aspect_ratio(first_frame, (1, 2.5), (2.5, 1))
-        image_list: list[OmniParamImage] = [
-            OmniParamImage(
-                image_url=(await upload_images_to_comfyapi(cls, first_frame, wait_label="Uploading first frame"))[0],
-                type="first_frame",
-            )
-        ]
-        if end_frame is not None:
-            validate_image_dimensions(end_frame, min_width=300, min_height=300)
-            validate_image_aspect_ratio(end_frame, (1, 2.5), (2.5, 1))
-            image_list.append(
-                OmniParamImage(
-                    image_url=(await upload_images_to_comfyapi(cls, end_frame, wait_label="Uploading end frame"))[0],
-                    type="end_frame",
-                )
-            )
-        if reference_images is not None:
-            if get_number_of_images(reference_images) > 6:
-                raise ValueError("The maximum number of reference images allowed is 6.")
-            for i in reference_images:
-                validate_image_dimensions(i, min_width=300, min_height=300)
-                validate_image_aspect_ratio(i, (1, 2.5), (2.5, 1))
-            for i in await upload_images_to_comfyapi(cls, reference_images, wait_label="Uploading reference frame(s)"):
-                image_list.append(OmniParamImage(image_url=i))
-        response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/kling/v1/videos/omni-video", method="POST"),
-            response_model=TaskStatusVideoResponse,
-            data=OmniProFirstLastFrameRequest(
-                model_name=model_name,
-                prompt=prompt,
-                duration=str(duration),
-                image_list=image_list,
-            ),
-        )
-        return await finish_omni_video_task(cls, response)
-
-
-class OmniProImageToVideoNode(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls) -> IO.Schema:
-        return IO.Schema(
-            node_id="KlingOmniProImageToVideoNode",
-            display_name="Kling Omni Image to Video (Pro)",
-            category="api node/video/Kling",
-            description="Use up to 7 reference images to generate a video with the latest Kling model.",
-            inputs=[
-                IO.Combo.Input("model_name", options=["kling-video-o1"]),
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    tooltip="A text prompt describing the video content. "
-                    "This can include both positive and negative descriptions.",
-                ),
-                IO.Combo.Input("aspect_ratio", options=["16:9", "9:16", "1:1"]),
-                IO.Int.Input("duration", default=3, min=3, max=10, display_mode=IO.NumberDisplay.slider),
-                IO.Image.Input(
-                    "reference_images",
-                    tooltip="Up to 7 reference images.",
-                ),
-            ],
-            outputs=[
-                IO.Video.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        model_name: str,
-        prompt: str,
-        aspect_ratio: str,
-        duration: int,
-        reference_images: Input.Image,
-    ) -> IO.NodeOutput:
-        validate_string(prompt, min_length=1, max_length=2500)
-        if get_number_of_images(reference_images) > 7:
-            raise ValueError("The maximum number of reference images is 7.")
-        for i in reference_images:
-            validate_image_dimensions(i, min_width=300, min_height=300)
-            validate_image_aspect_ratio(i, (1, 2.5), (2.5, 1))
-        image_list: list[OmniParamImage] = []
-        for i in await upload_images_to_comfyapi(cls, reference_images, wait_label="Uploading reference image"):
-            image_list.append(OmniParamImage(image_url=i))
-        response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/kling/v1/videos/omni-video", method="POST"),
-            response_model=TaskStatusVideoResponse,
-            data=OmniProReferences2VideoRequest(
-                model_name=model_name,
-                prompt=prompt,
-                aspect_ratio=aspect_ratio,
-                duration=str(duration),
-                image_list=image_list,
-            ),
-        )
-        return await finish_omni_video_task(cls, response)
-
-
-class OmniProVideoToVideoNode(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls) -> IO.Schema:
-        return IO.Schema(
-            node_id="KlingOmniProVideoToVideoNode",
-            display_name="Kling Omni Video to Video (Pro)",
-            category="api node/video/Kling",
-            description="Use a video and up to 4 reference images to generate a video with the latest Kling model.",
-            inputs=[
-                IO.Combo.Input("model_name", options=["kling-video-o1"]),
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    tooltip="A text prompt describing the video content. "
-                    "This can include both positive and negative descriptions.",
-                ),
-                IO.Combo.Input("aspect_ratio", options=["16:9", "9:16", "1:1"]),
-                IO.Int.Input("duration", default=3, min=3, max=10, display_mode=IO.NumberDisplay.slider),
-                IO.Video.Input("reference_video", tooltip="Video to use as a reference."),
-                IO.Boolean.Input("keep_original_sound", default=True),
-                IO.Image.Input(
-                    "reference_images",
-                    tooltip="Up to 4 additional reference images.",
-                    optional=True,
-                ),
-            ],
-            outputs=[
-                IO.Video.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        model_name: str,
-        prompt: str,
-        aspect_ratio: str,
-        duration: int,
-        reference_video: Input.Video,
-        keep_original_sound: bool,
-        reference_images: Input.Image | None = None,
-    ) -> IO.NodeOutput:
-        validate_string(prompt, min_length=1, max_length=2500)
-        validate_video_duration(reference_video, min_duration=3.0, max_duration=10.05)
-        validate_video_dimensions(reference_video, min_width=720, min_height=720, max_width=2160, max_height=2160)
-        image_list: list[OmniParamImage] = []
-        if reference_images is not None:
-            if get_number_of_images(reference_images) > 4:
-                raise ValueError("The maximum number of reference images allowed with a video input is 4.")
-            for i in reference_images:
-                validate_image_dimensions(i, min_width=300, min_height=300)
-                validate_image_aspect_ratio(i, (1, 2.5), (2.5, 1))
-            for i in await upload_images_to_comfyapi(cls, reference_images, wait_label="Uploading reference image"):
-                image_list.append(OmniParamImage(image_url=i))
-        video_list = [
-            OmniParamVideo(
-                video_url=await upload_video_to_comfyapi(cls, reference_video, wait_label="Uploading reference video"),
-                refer_type="feature",
-                keep_original_sound="yes" if keep_original_sound else "no",
-            )
-        ]
-        response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/kling/v1/videos/omni-video", method="POST"),
-            response_model=TaskStatusVideoResponse,
-            data=OmniProReferences2VideoRequest(
-                model_name=model_name,
-                prompt=prompt,
-                aspect_ratio=aspect_ratio,
-                duration=str(duration),
-                image_list=image_list if image_list else None,
-                video_list=video_list,
-            ),
-        )
-        return await finish_omni_video_task(cls, response)
-
-
-class OmniProEditVideoNode(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls) -> IO.Schema:
-        return IO.Schema(
-            node_id="KlingOmniProEditVideoNode",
-            display_name="Kling Omni Edit Video (Pro)",
-            category="api node/video/Kling",
-            description="Edit an existing video with the latest model from Kling.",
-            inputs=[
-                IO.Combo.Input("model_name", options=["kling-video-o1"]),
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    tooltip="A text prompt describing the video content. "
-                    "This can include both positive and negative descriptions.",
-                ),
-                IO.Video.Input("video", tooltip="Video for editing. The output video length will be the same."),
-                IO.Boolean.Input("keep_original_sound", default=True),
-                IO.Image.Input(
-                    "reference_images",
-                    tooltip="Up to 4 additional reference images.",
-                    optional=True,
-                ),
-            ],
-            outputs=[
-                IO.Video.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        model_name: str,
-        prompt: str,
-        video: Input.Video,
-        keep_original_sound: bool,
-        reference_images: Input.Image | None = None,
-    ) -> IO.NodeOutput:
-        validate_string(prompt, min_length=1, max_length=2500)
-        validate_video_duration(video, min_duration=3.0, max_duration=10.05)
-        validate_video_dimensions(video, min_width=720, min_height=720, max_width=2160, max_height=2160)
-        image_list: list[OmniParamImage] = []
-        if reference_images is not None:
-            if get_number_of_images(reference_images) > 4:
-                raise ValueError("The maximum number of reference images allowed with a video input is 4.")
-            for i in reference_images:
-                validate_image_dimensions(i, min_width=300, min_height=300)
-                validate_image_aspect_ratio(i, (1, 2.5), (2.5, 1))
-            for i in await upload_images_to_comfyapi(cls, reference_images, wait_label="Uploading reference image"):
-                image_list.append(OmniParamImage(image_url=i))
-        video_list = [
-            OmniParamVideo(
-                video_url=await upload_video_to_comfyapi(cls, video, wait_label="Uploading base video"),
-                refer_type="base",
-                keep_original_sound="yes" if keep_original_sound else "no",
-            )
-        ]
-        response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/kling/v1/videos/omni-video", method="POST"),
-            response_model=TaskStatusVideoResponse,
-            data=OmniProReferences2VideoRequest(
-                model_name=model_name,
-                prompt=prompt,
-                aspect_ratio=None,
-                duration=None,
-                image_list=image_list if image_list else None,
-                video_list=video_list,
-            ),
-        )
-        return await finish_omni_video_task(cls, response)
-
-
 class KlingCameraControlT2VNode(IO.ComfyNode):
    """
    Kling Text to Video Camera Control Node. This node is a text to video node, but it supports controlling the camera.
@@ -1185,7 +785,7 @@ class KlingCameraControlT2VNode(IO.ComfyNode):
        negative_prompt: str,
        cfg_scale: float,
        aspect_ratio: str,
-        camera_control: KlingCameraControl | None = None,
+        camera_control: Optional[KlingCameraControl] = None,
    ) -> IO.NodeOutput:
        return await execute_text2video(
            cls,
@@ -1252,8 +852,8 @@ class KlingImage2VideoNode(IO.ComfyNode):
        mode: str,
        aspect_ratio: str,
        duration: str,
-        camera_control: KlingCameraControl | None = None,
-        end_frame: torch.Tensor | None = None,
+        camera_control: Optional[KlingCameraControl] = None,
+        end_frame: Optional[torch.Tensor] = None,
    ) -> IO.NodeOutput:
        return await execute_image2video(
            cls,
@@ -1363,11 +963,15 @@ class KlingStartEndFrameNode(IO.ComfyNode):
                IO.String.Input("prompt", multiline=True, tooltip="Positive text prompt"),
                IO.String.Input("negative_prompt", multiline=True, tooltip="Negative text prompt"),
                IO.Float.Input("cfg_scale", default=0.5, min=0.0, max=1.0),
-                IO.Combo.Input("aspect_ratio", options=["16:9", "9:16", "1:1"]),
+                IO.Combo.Input(
+                    "aspect_ratio",
+                    options=[i.value for i in KlingVideoGenAspectRatio],
+                    default="16:9",
+                ),
                IO.Combo.Input(
                    "mode",
                    options=modes,
-                    default=modes[8],
+                    default=modes[2],
                    tooltip="The configuration to use for the video generation following the format: mode / duration / model_name.",
                ),
            ],
@@ -1564,10 +1168,7 @@ class KlingSingleImageVideoEffectNode(IO.ComfyNode):
            category="api node/video/Kling",
            description="Achieve different special effects when generating a video based on the effect_scene.",
            inputs=[
-                IO.Image.Input(
-                    "image",
-                    tooltip=" Reference Image. URL or Base64 encoded string (without data:image prefix). File size cannot exceed 10MB, resolution not less than 300*300px, aspect ratio between 1:2.5 ~ 2.5:1",
-                ),
+                IO.Image.Input("image", tooltip=" Reference Image. URL or Base64 encoded string (without data:image prefix). File size cannot exceed 10MB, resolution not less than 300*300px, aspect ratio between 1:2.5 ~ 2.5:1"),
                IO.Combo.Input(
                    "effect_scene",
                    options=[i.value for i in KlingSingleImageEffectsScene],
@@ -1651,8 +1252,8 @@ class KlingLipSyncAudioToVideoNode(IO.ComfyNode):
    @classmethod
    async def execute(
        cls,
-        video: Input.Video,
-        audio: Input.Audio,
+        video: VideoInput,
+        audio: AudioInput,
        voice_language: str,
    ) -> IO.NodeOutput:
        return await execute_lipsync(
@@ -1711,7 +1312,7 @@ class KlingLipSyncTextToVideoNode(IO.ComfyNode):
    @classmethod
    async def execute(
        cls,
-        video: Input.Video,
+        video: VideoInput,
        text: str,
        voice: str,
        voice_speed: float,
@@ -1868,7 +1469,7 @@ class KlingImageGenerationNode(IO.ComfyNode):
        human_fidelity: float,
        n: int,
        aspect_ratio: KlingImageGenAspectRatio,
-        image: torch.Tensor | None = None,
+        image: Optional[torch.Tensor] = None,
    ) -> IO.NodeOutput:
        validate_string(prompt, field_name="prompt", min_length=1, max_length=MAX_PROMPT_LENGTH_IMAGE_GEN)
        validate_string(negative_prompt, field_name="negative_prompt", max_length=MAX_PROMPT_LENGTH_IMAGE_GEN)
@@ -1930,11 +1531,6 @@ class KlingExtension(ComfyExtension):
            KlingImageGenerationNode,
            KlingSingleImageVideoEffectNode,
            KlingDualCharacterVideoEffectNode,
-            OmniProTextToVideoNode,
-            OmniProFirstLastFrameNode,
-            OmniProImageToVideoNode,
-            OmniProVideoToVideoNode,
-            OmniProEditVideoNode,
        ]


--- a/comfy_api_nodes/nodes_ltxv.py
+++ b/comfy_api_nodes/nodes_ltxv.py
@@ -46,7 +46,7 @@ class TextToVideoNode(IO.ComfyNode):
                    multiline=True,
                    default="",
                ),
-                IO.Combo.Input("duration", options=[6, 8, 10, 12, 14, 16, 18, 20], default=8),
+                IO.Combo.Input("duration", options=[6, 8, 10], default=8),
                IO.Combo.Input(
                    "resolution",
                    options=[
@@ -85,10 +85,6 @@ class TextToVideoNode(IO.ComfyNode):
        generate_audio: bool = False,
    ) -> IO.NodeOutput:
        validate_string(prompt, min_length=1, max_length=10000)
-        if duration > 10 and (model != "LTX-2 (Fast)" or resolution != "1920x1080" or fps != 25):
-            raise ValueError(
-                "Durations over 10s are only available for the Fast model at 1920x1080 resolution and 25 FPS."
-            )
        response = await sync_op_raw(
            cls,
            ApiEndpoint("/proxy/ltx/v1/text-to-video", "POST"),
@@ -122,7 +118,7 @@ class ImageToVideoNode(IO.ComfyNode):
                    multiline=True,
                    default="",
                ),
-                IO.Combo.Input("duration", options=[6, 8, 10, 12, 14, 16, 18, 20], default=8),
+                IO.Combo.Input("duration", options=[6, 8, 10], default=8),
                IO.Combo.Input(
                    "resolution",
                    options=[
@@ -162,10 +158,6 @@ class ImageToVideoNode(IO.ComfyNode):
        generate_audio: bool = False,
    ) -> IO.NodeOutput:
        validate_string(prompt, min_length=1, max_length=10000)
-        if duration > 10 and (model != "LTX-2 (Fast)" or resolution != "1920x1080" or fps != 25):
-            raise ValueError(
-                "Durations over 10s are only available for the Fast model at 1920x1080 resolution and 25 FPS."
-            )
        if get_number_of_images(image) != 1:
            raise ValueError("Currently only one input image is supported.")
        response = await sync_op_raw(
--- a/comfy_api_nodes/nodes_luma.py
+++ b/comfy_api_nodes/nodes_luma.py
@@ -1,51 +1,69 @@
+from __future__ import annotations
+from inspect import cleandoc
 from typing import Optional
-
-import torch
 from typing_extensions import override
-
-from comfy_api.latest import IO, ComfyExtension
+from comfy_api.latest import ComfyExtension, IO
+from comfy_api.input_impl.video_types import VideoFromFile
 from comfy_api_nodes.apis.luma_api import (
-    LumaAspectRatio,
-    LumaCharacterRef,
-    LumaConceptChain,
-    LumaGeneration,
-    LumaGenerationRequest,
-    LumaImageGenerationRequest,
-    LumaImageIdentity,
    LumaImageModel,
-    LumaImageReference,
-    LumaIO,
-    LumaKeyframes,
+    LumaVideoModel,
+    LumaVideoOutputResolution,
+    LumaVideoModelOutputDuration,
+    LumaAspectRatio,
+    LumaState,
+    LumaImageGenerationRequest,
+    LumaGenerationRequest,
+    LumaGeneration,
+    LumaCharacterRef,
    LumaModifyImageRef,
+    LumaImageIdentity,
    LumaReference,
    LumaReferenceChain,
-    LumaVideoModel,
-    LumaVideoModelOutputDuration,
-    LumaVideoOutputResolution,
+    LumaImageReference,
+    LumaKeyframes,
+    LumaConceptChain,
+    LumaIO,
    get_luma_concepts,
 )
-from comfy_api_nodes.util import (
+from comfy_api_nodes.apis.client import (
    ApiEndpoint,
-    download_url_to_image_tensor,
-    download_url_to_video_output,
-    poll_op,
-    sync_op,
-    upload_images_to_comfyapi,
-    validate_string,
+    HttpMethod,
+    SynchronousOperation,
+    PollingOperation,
+    EmptyRequest,
 )
+from comfy_api_nodes.apinode_utils import (
+    upload_images_to_comfyapi,
+    process_image_response,
+)
+from server import PromptServer
+from comfy_api_nodes.util import validate_string
+
+import aiohttp
+import torch
+from io import BytesIO

 LUMA_T2V_AVERAGE_DURATION = 105
 LUMA_I2V_AVERAGE_DURATION = 100

+def image_result_url_extractor(response: LumaGeneration):
+    return response.assets.image if hasattr(response, "assets") and hasattr(response.assets, "image") else None
+
+def video_result_url_extractor(response: LumaGeneration):
+    return response.assets.video if hasattr(response, "assets") and hasattr(response.assets, "video") else None

 class LumaReferenceNode(IO.ComfyNode):
+    """
+    Holds an image and weight for use with Luma Generate Image node.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="LumaReferenceNode",
            display_name="Luma Reference",
            category="api node/image/Luma",
-            description="Holds an image and weight for use with Luma Generate Image node.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.Image.Input(
                    "image",
@@ -65,10 +83,17 @@ class LumaReferenceNode(IO.ComfyNode):
                ),
            ],
            outputs=[IO.Custom(LumaIO.LUMA_REF).Output(display_name="luma_ref")],
+            hidden=[
+                IO.Hidden.auth_token_comfy_org,
+                IO.Hidden.api_key_comfy_org,
+                IO.Hidden.unique_id,
+            ],
        )

    @classmethod
-    def execute(cls, image: torch.Tensor, weight: float, luma_ref: LumaReferenceChain = None) -> IO.NodeOutput:
+    def execute(
+        cls, image: torch.Tensor, weight: float, luma_ref: LumaReferenceChain = None
+    ) -> IO.NodeOutput:
        if luma_ref is not None:
            luma_ref = luma_ref.clone()
        else:
@@ -78,13 +103,17 @@ class LumaReferenceNode(IO.ComfyNode):


 class LumaConceptsNode(IO.ComfyNode):
+    """
+    Holds one or more Camera Concepts for use with Luma Text to Video and Luma Image to Video nodes.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="LumaConceptsNode",
            display_name="Luma Concepts",
            category="api node/video/Luma",
-            description="Camera Concepts for use with Luma Text to Video and Luma Image to Video nodes.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.Combo.Input(
                    "concept1",
@@ -109,6 +138,11 @@ class LumaConceptsNode(IO.ComfyNode):
                ),
            ],
            outputs=[IO.Custom(LumaIO.LUMA_CONCEPTS).Output(display_name="luma_concepts")],
+            hidden=[
+                IO.Hidden.auth_token_comfy_org,
+                IO.Hidden.api_key_comfy_org,
+                IO.Hidden.unique_id,
+            ],
        )

    @classmethod
@@ -127,13 +161,17 @@ class LumaConceptsNode(IO.ComfyNode):


 class LumaImageGenerationNode(IO.ComfyNode):
+    """
+    Generates images synchronously based on prompt and aspect ratio.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="LumaImageNode",
            display_name="Luma Text to Image",
            category="api node/image/Luma",
-            description="Generates images synchronously based on prompt and aspect ratio.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.String.Input(
                    "prompt",
@@ -199,30 +237,45 @@ class LumaImageGenerationNode(IO.ComfyNode):
        aspect_ratio: str,
        seed,
        style_image_weight: float,
-        image_luma_ref: Optional[LumaReferenceChain] = None,
-        style_image: Optional[torch.Tensor] = None,
-        character_image: Optional[torch.Tensor] = None,
+        image_luma_ref: LumaReferenceChain = None,
+        style_image: torch.Tensor = None,
+        character_image: torch.Tensor = None,
    ) -> IO.NodeOutput:
        validate_string(prompt, strip_whitespace=True, min_length=3)
+        auth_kwargs = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        # handle image_luma_ref
        api_image_ref = None
        if image_luma_ref is not None:
-            api_image_ref = await cls._convert_luma_refs(image_luma_ref, max_refs=4)
+            api_image_ref = await cls._convert_luma_refs(
+                image_luma_ref, max_refs=4, auth_kwargs=auth_kwargs,
+            )
        # handle style_luma_ref
        api_style_ref = None
        if style_image is not None:
-            api_style_ref = await cls._convert_style_image(style_image, weight=style_image_weight)
+            api_style_ref = await cls._convert_style_image(
+                style_image, weight=style_image_weight, auth_kwargs=auth_kwargs,
+            )
        # handle character_ref images
        character_ref = None
        if character_image is not None:
-            download_urls = await upload_images_to_comfyapi(cls, character_image, max_images=4)
-            character_ref = LumaCharacterRef(identity0=LumaImageIdentity(images=download_urls))
+            download_urls = await upload_images_to_comfyapi(
+                character_image, max_images=4, auth_kwargs=auth_kwargs,
+            )
+            character_ref = LumaCharacterRef(
+                identity0=LumaImageIdentity(images=download_urls)
+            )

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/luma/generations/image", method="POST"),
-            response_model=LumaGeneration,
-            data=LumaImageGenerationRequest(
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/luma/generations/image",
+                method=HttpMethod.POST,
+                request_model=LumaImageGenerationRequest,
+                response_model=LumaGeneration,
+            ),
+            request=LumaImageGenerationRequest(
                prompt=prompt,
                model=model,
                aspect_ratio=aspect_ratio,
@@ -230,21 +283,41 @@ class LumaImageGenerationNode(IO.ComfyNode):
                style_ref=api_style_ref,
                character_ref=character_ref,
            ),
+            auth_kwargs=auth_kwargs,
        )
-        response_poll = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/luma/generations/{response_api.id}"),
-            response_model=LumaGeneration,
+        response_api: LumaGeneration = await operation.execute()
+
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/luma/generations/{response_api.id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=LumaGeneration,
+            ),
+            completed_statuses=[LumaState.completed],
+            failed_statuses=[LumaState.failed],
            status_extractor=lambda x: x.state,
+            result_url_extractor=image_result_url_extractor,
+            node_id=cls.hidden.unique_id,
+            auth_kwargs=auth_kwargs,
        )
-        return IO.NodeOutput(await download_url_to_image_tensor(response_poll.assets.image))
+        response_poll = await operation.execute()
+
+        async with aiohttp.ClientSession() as session:
+            async with session.get(response_poll.assets.image) as img_response:
+                img = process_image_response(await img_response.content.read())
+        return IO.NodeOutput(img)

    @classmethod
-    async def _convert_luma_refs(cls, luma_ref: LumaReferenceChain, max_refs: int):
+    async def _convert_luma_refs(
+        cls, luma_ref: LumaReferenceChain, max_refs: int, auth_kwargs: Optional[dict[str,str]] = None
+    ):
        luma_urls = []
        ref_count = 0
        for ref in luma_ref.refs:
-            download_urls = await upload_images_to_comfyapi(cls, ref.image, max_images=1)
+            download_urls = await upload_images_to_comfyapi(
+                ref.image, max_images=1, auth_kwargs=auth_kwargs
+            )
            luma_urls.append(download_urls[0])
            ref_count += 1
            if ref_count >= max_refs:
@@ -252,19 +325,27 @@ class LumaImageGenerationNode(IO.ComfyNode):
        return luma_ref.create_api_model(download_urls=luma_urls, max_refs=max_refs)

    @classmethod
-    async def _convert_style_image(cls, style_image: torch.Tensor, weight: float):
-        chain = LumaReferenceChain(first_ref=LumaReference(image=style_image, weight=weight))
-        return await cls._convert_luma_refs(chain, max_refs=1)
+    async def _convert_style_image(
+        cls, style_image: torch.Tensor, weight: float, auth_kwargs: Optional[dict[str,str]] = None
+    ):
+        chain = LumaReferenceChain(
+            first_ref=LumaReference(image=style_image, weight=weight)
+        )
+        return await cls._convert_luma_refs(chain, max_refs=1, auth_kwargs=auth_kwargs)


 class LumaImageModifyNode(IO.ComfyNode):
+    """
+    Modifies images synchronously based on prompt and aspect ratio.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="LumaImageModifyNode",
            display_name="Luma Image to Image",
            category="api node/image/Luma",
-            description="Modifies images synchronously based on prompt and aspect ratio.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.Image.Input(
                    "image",
@@ -314,37 +395,68 @@ class LumaImageModifyNode(IO.ComfyNode):
        image_weight: float,
        seed,
    ) -> IO.NodeOutput:
-        download_urls = await upload_images_to_comfyapi(cls, image, max_images=1)
+        auth_kwargs = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        # first, upload image
+        download_urls = await upload_images_to_comfyapi(
+            image, max_images=1, auth_kwargs=auth_kwargs,
+        )
        image_url = download_urls[0]
-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/luma/generations/image", method="POST"),
-            response_model=LumaGeneration,
-            data=LumaImageGenerationRequest(
+        # next, make Luma call with download url provided
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/luma/generations/image",
+                method=HttpMethod.POST,
+                request_model=LumaImageGenerationRequest,
+                response_model=LumaGeneration,
+            ),
+            request=LumaImageGenerationRequest(
                prompt=prompt,
                model=model,
                modify_image_ref=LumaModifyImageRef(
-                    url=image_url, weight=round(max(min(1.0 - image_weight, 0.98), 0.0), 2)
+                    url=image_url, weight=round(max(min(1.0-image_weight, 0.98), 0.0), 2)
                ),
            ),
+            auth_kwargs=auth_kwargs,
        )
-        response_poll = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/luma/generations/{response_api.id}"),
-            response_model=LumaGeneration,
+        response_api: LumaGeneration = await operation.execute()
+
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/luma/generations/{response_api.id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=LumaGeneration,
+            ),
+            completed_statuses=[LumaState.completed],
+            failed_statuses=[LumaState.failed],
            status_extractor=lambda x: x.state,
+            result_url_extractor=image_result_url_extractor,
+            node_id=cls.hidden.unique_id,
+            auth_kwargs=auth_kwargs,
        )
-        return IO.NodeOutput(await download_url_to_image_tensor(response_poll.assets.image))
+        response_poll = await operation.execute()
+
+        async with aiohttp.ClientSession() as session:
+            async with session.get(response_poll.assets.image) as img_response:
+                img = process_image_response(await img_response.content.read())
+        return IO.NodeOutput(img)


 class LumaTextToVideoGenerationNode(IO.ComfyNode):
+    """
+    Generates videos synchronously based on prompt and output_size.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="LumaVideoNode",
            display_name="Luma Text to Video",
            category="api node/video/Luma",
-            description="Generates videos synchronously based on prompt and output_size.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.String.Input(
                    "prompt",
@@ -386,7 +498,7 @@ class LumaTextToVideoGenerationNode(IO.ComfyNode):
                    "luma_concepts",
                    tooltip="Optional Camera Concepts to dictate camera motion via the Luma Concepts node.",
                    optional=True,
-                ),
+                )
            ],
            outputs=[IO.Video.Output()],
            hidden=[
@@ -407,17 +519,24 @@ class LumaTextToVideoGenerationNode(IO.ComfyNode):
        duration: str,
        loop: bool,
        seed,
-        luma_concepts: Optional[LumaConceptChain] = None,
+        luma_concepts: LumaConceptChain = None,
    ) -> IO.NodeOutput:
        validate_string(prompt, strip_whitespace=False, min_length=3)
        duration = duration if model != LumaVideoModel.ray_1_6 else None
        resolution = resolution if model != LumaVideoModel.ray_1_6 else None

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/luma/generations", method="POST"),
-            response_model=LumaGeneration,
-            data=LumaGenerationRequest(
+        auth_kwargs = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/luma/generations",
+                method=HttpMethod.POST,
+                request_model=LumaGenerationRequest,
+                response_model=LumaGeneration,
+            ),
+            request=LumaGenerationRequest(
                prompt=prompt,
                model=model,
                resolution=resolution,
@@ -426,25 +545,47 @@ class LumaTextToVideoGenerationNode(IO.ComfyNode):
                loop=loop,
                concepts=luma_concepts.create_api_model() if luma_concepts else None,
            ),
+            auth_kwargs=auth_kwargs,
        )
-        response_poll = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/luma/generations/{response_api.id}"),
-            response_model=LumaGeneration,
+        response_api: LumaGeneration = await operation.execute()
+
+        if cls.hidden.unique_id:
+            PromptServer.instance.send_progress_text(f"Luma video generation started: {response_api.id}", cls.hidden.unique_id)
+
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/luma/generations/{response_api.id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=LumaGeneration,
+            ),
+            completed_statuses=[LumaState.completed],
+            failed_statuses=[LumaState.failed],
            status_extractor=lambda x: x.state,
+            result_url_extractor=video_result_url_extractor,
+            node_id=cls.hidden.unique_id,
            estimated_duration=LUMA_T2V_AVERAGE_DURATION,
+            auth_kwargs=auth_kwargs,
        )
-        return IO.NodeOutput(await download_url_to_video_output(response_poll.assets.video))
+        response_poll = await operation.execute()
+
+        async with aiohttp.ClientSession() as session:
+            async with session.get(response_poll.assets.video) as vid_response:
+                return IO.NodeOutput(VideoFromFile(BytesIO(await vid_response.content.read())))


 class LumaImageToVideoGenerationNode(IO.ComfyNode):
+    """
+    Generates videos synchronously based on prompt, input images, and output_size.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="LumaImageToVideoNode",
            display_name="Luma Image to Video",
            category="api node/video/Luma",
-            description="Generates videos synchronously based on prompt, input images, and output_size.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.String.Input(
                    "prompt",
@@ -496,7 +637,7 @@ class LumaImageToVideoGenerationNode(IO.ComfyNode):
                    "luma_concepts",
                    tooltip="Optional Camera Concepts to dictate camera motion via the Luma Concepts node.",
                    optional=True,
-                ),
+                )
            ],
            outputs=[IO.Video.Output()],
            hidden=[
@@ -521,15 +662,25 @@ class LumaImageToVideoGenerationNode(IO.ComfyNode):
        luma_concepts: LumaConceptChain = None,
    ) -> IO.NodeOutput:
        if first_image is None and last_image is None:
-            raise Exception("At least one of first_image and last_image requires an input.")
-        keyframes = await cls._convert_to_keyframes(first_image, last_image)
+            raise Exception(
+                "At least one of first_image and last_image requires an input."
+            )
+        auth_kwargs = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        keyframes = await cls._convert_to_keyframes(first_image, last_image, auth_kwargs=auth_kwargs)
        duration = duration if model != LumaVideoModel.ray_1_6 else None
        resolution = resolution if model != LumaVideoModel.ray_1_6 else None
-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/luma/generations", method="POST"),
-            response_model=LumaGeneration,
-            data=LumaGenerationRequest(
+
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/luma/generations",
+                method=HttpMethod.POST,
+                request_model=LumaGenerationRequest,
+                response_model=LumaGeneration,
+            ),
+            request=LumaGenerationRequest(
                prompt=prompt,
                model=model,
                aspect_ratio=LumaAspectRatio.ratio_16_9,  # ignored, but still needed by the API for some reason
@@ -539,31 +690,54 @@ class LumaImageToVideoGenerationNode(IO.ComfyNode):
                keyframes=keyframes,
                concepts=luma_concepts.create_api_model() if luma_concepts else None,
            ),
+            auth_kwargs=auth_kwargs,
        )
-        response_poll = await poll_op(
-            cls,
-            poll_endpoint=ApiEndpoint(path=f"/proxy/luma/generations/{response_api.id}"),
-            response_model=LumaGeneration,
+        response_api: LumaGeneration = await operation.execute()
+
+        if cls.hidden.unique_id:
+            PromptServer.instance.send_progress_text(f"Luma video generation started: {response_api.id}", cls.hidden.unique_id)
+
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/luma/generations/{response_api.id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=LumaGeneration,
+            ),
+            completed_statuses=[LumaState.completed],
+            failed_statuses=[LumaState.failed],
            status_extractor=lambda x: x.state,
+            result_url_extractor=video_result_url_extractor,
+            node_id=cls.hidden.unique_id,
            estimated_duration=LUMA_I2V_AVERAGE_DURATION,
+            auth_kwargs=auth_kwargs,
        )
-        return IO.NodeOutput(await download_url_to_video_output(response_poll.assets.video))
+        response_poll = await operation.execute()
+
+        async with aiohttp.ClientSession() as session:
+            async with session.get(response_poll.assets.video) as vid_response:
+                return IO.NodeOutput(VideoFromFile(BytesIO(await vid_response.content.read())))

    @classmethod
    async def _convert_to_keyframes(
        cls,
        first_image: torch.Tensor = None,
        last_image: torch.Tensor = None,
+        auth_kwargs: Optional[dict[str,str]] = None,
    ):
        if first_image is None and last_image is None:
            return None
        frame0 = None
        frame1 = None
        if first_image is not None:
-            download_urls = await upload_images_to_comfyapi(cls, first_image, max_images=1)
+            download_urls = await upload_images_to_comfyapi(
+                first_image, max_images=1, auth_kwargs=auth_kwargs,
+            )
            frame0 = LumaImageReference(type="image", url=download_urls[0])
        if last_image is not None:
-            download_urls = await upload_images_to_comfyapi(cls, last_image, max_images=1)
+            download_urls = await upload_images_to_comfyapi(
+                last_image, max_images=1, auth_kwargs=auth_kwargs,
+            )
            frame1 = LumaImageReference(type="image", url=download_urls[0])
        return LumaKeyframes(frame0=frame0, frame1=frame1)

--- a/comfy_api_nodes/nodes_minimax.py
+++ b/comfy_api_nodes/nodes_minimax.py
@@ -1,57 +1,71 @@
+from inspect import cleandoc
 from typing import Optional
-
+import logging
 import torch
-from typing_extensions import override

-from comfy_api.latest import IO, ComfyExtension
-from comfy_api_nodes.apis.minimax_api import (
-    MinimaxFileRetrieveResponse,
-    MiniMaxModel,
-    MinimaxTaskResultResponse,
+from typing_extensions import override
+from comfy_api.latest import ComfyExtension, IO
+from comfy_api.input_impl.video_types import VideoFromFile
+from comfy_api_nodes.apis import (
    MinimaxVideoGenerationRequest,
    MinimaxVideoGenerationResponse,
+    MinimaxFileRetrieveResponse,
+    MinimaxTaskResultResponse,
    SubjectReferenceItem,
+    MiniMaxModel,
 )
-from comfy_api_nodes.util import (
+from comfy_api_nodes.apis.client import (
    ApiEndpoint,
-    download_url_to_video_output,
-    poll_op,
-    sync_op,
-    upload_images_to_comfyapi,
-    validate_string,
+    HttpMethod,
+    SynchronousOperation,
+    PollingOperation,
+    EmptyRequest,
 )
+from comfy_api_nodes.apinode_utils import (
+    download_url_to_bytesio,
+    upload_images_to_comfyapi,
+)
+from comfy_api_nodes.util import validate_string
+from server import PromptServer
+

 I2V_AVERAGE_DURATION = 114
 T2V_AVERAGE_DURATION = 234


 async def _generate_mm_video(
-    cls: type[IO.ComfyNode],
    *,
+    auth: dict[str, str],
+    node_id: str,
    prompt_text: str,
    seed: int,
    model: str,
-    image: Optional[torch.Tensor] = None,  # used for ImageToVideo
-    subject: Optional[torch.Tensor] = None,  # used for SubjectToVideo
+    image: Optional[torch.Tensor] = None,   # used for ImageToVideo
+    subject: Optional[torch.Tensor] = None, # used for SubjectToVideo
    average_duration: Optional[int] = None,
 ) -> IO.NodeOutput:
    if image is None:
        validate_string(prompt_text, field_name="prompt_text")
+    # upload image, if passed in
    image_url = None
    if image is not None:
-        image_url = (await upload_images_to_comfyapi(cls, image, max_images=1))[0]
+        image_url = (await upload_images_to_comfyapi(image, max_images=1, auth_kwargs=auth))[0]

    # TODO: figure out how to deal with subject properly, API returns invalid params when using S2V-01 model
    subject_reference = None
    if subject is not None:
-        subject_url = (await upload_images_to_comfyapi(cls, subject, max_images=1))[0]
+        subject_url = (await upload_images_to_comfyapi(subject, max_images=1, auth_kwargs=auth))[0]
        subject_reference = [SubjectReferenceItem(image=subject_url)]

-    response = await sync_op(
-        cls,
-        ApiEndpoint(path="/proxy/minimax/video_generation", method="POST"),
-        response_model=MinimaxVideoGenerationResponse,
-        data=MinimaxVideoGenerationRequest(
+
+    video_generate_operation = SynchronousOperation(
+        endpoint=ApiEndpoint(
+            path="/proxy/minimax/video_generation",
+            method=HttpMethod.POST,
+            request_model=MinimaxVideoGenerationRequest,
+            response_model=MinimaxVideoGenerationResponse,
+        ),
+        request=MinimaxVideoGenerationRequest(
            model=MiniMaxModel(model),
            prompt=prompt_text,
            callback_url=None,
@@ -59,50 +73,81 @@ async def _generate_mm_video(
            subject_reference=subject_reference,
            prompt_optimizer=None,
        ),
+        auth_kwargs=auth,
    )
+    response = await video_generate_operation.execute()

    task_id = response.task_id
    if not task_id:
        raise Exception(f"MiniMax generation failed: {response.base_resp}")

-    task_result = await poll_op(
-        cls,
-        ApiEndpoint(path="/proxy/minimax/query/video_generation", query_params={"task_id": task_id}),
-        response_model=MinimaxTaskResultResponse,
+    video_generate_operation = PollingOperation(
+        poll_endpoint=ApiEndpoint(
+            path="/proxy/minimax/query/video_generation",
+            method=HttpMethod.GET,
+            request_model=EmptyRequest,
+            response_model=MinimaxTaskResultResponse,
+            query_params={"task_id": task_id},
+        ),
+        completed_statuses=["Success"],
+        failed_statuses=["Fail"],
        status_extractor=lambda x: x.status.value,
        estimated_duration=average_duration,
+        node_id=node_id,
+        auth_kwargs=auth,
    )
+    task_result = await video_generate_operation.execute()

    file_id = task_result.file_id
    if file_id is None:
        raise Exception("Request was not successful. Missing file ID.")
-    file_result = await sync_op(
-        cls,
-        ApiEndpoint(path="/proxy/minimax/files/retrieve", query_params={"file_id": int(file_id)}),
-        response_model=MinimaxFileRetrieveResponse,
+    file_retrieve_operation = SynchronousOperation(
+        endpoint=ApiEndpoint(
+            path="/proxy/minimax/files/retrieve",
+            method=HttpMethod.GET,
+            request_model=EmptyRequest,
+            response_model=MinimaxFileRetrieveResponse,
+            query_params={"file_id": int(file_id)},
+        ),
+        request=EmptyRequest(),
+        auth_kwargs=auth,
    )
+    file_result = await file_retrieve_operation.execute()

    file_url = file_result.file.download_url
    if file_url is None:
-        raise Exception(f"No video was found in the response. Full response: {file_result.model_dump()}")
-    if file_result.file.backup_download_url:
-        try:
-            return IO.NodeOutput(await download_url_to_video_output(file_url, timeout=10, max_retries=2))
-        except Exception:  # if we have a second URL to retrieve the result, try again using that one
-            return IO.NodeOutput(
-                await download_url_to_video_output(file_result.file.backup_download_url, max_retries=3)
-            )
-    return IO.NodeOutput(await download_url_to_video_output(file_url))
+        raise Exception(
+            f"No video was found in the response. Full response: {file_result.model_dump()}"
+        )
+    logging.info("Generated video URL: %s", file_url)
+    if node_id:
+        if hasattr(file_result.file, "backup_download_url"):
+            message = f"Result URL: {file_url}\nBackup URL: {file_result.file.backup_download_url}"
+        else:
+            message = f"Result URL: {file_url}"
+        PromptServer.instance.send_progress_text(message, node_id)
+
+    # Download and return as VideoFromFile
+    video_io = await download_url_to_bytesio(file_url)
+    if video_io is None:
+        error_msg = f"Failed to download video from {file_url}"
+        logging.error(error_msg)
+        raise Exception(error_msg)
+    return IO.NodeOutput(VideoFromFile(video_io))


 class MinimaxTextToVideoNode(IO.ComfyNode):
+    """
+    Generates videos synchronously based on a prompt, and optional parameters using MiniMax's API.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="MinimaxTextToVideoNode",
            display_name="MiniMax Text to Video",
            category="api node/video/MiniMax",
-            description="Generates videos synchronously based on a prompt, and optional parameters.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.String.Input(
                    "prompt_text",
@@ -144,7 +189,11 @@ class MinimaxTextToVideoNode(IO.ComfyNode):
        seed: int = 0,
    ) -> IO.NodeOutput:
        return await _generate_mm_video(
-            cls,
+            auth={
+                "auth_token": cls.hidden.auth_token_comfy_org,
+                "comfy_api_key": cls.hidden.api_key_comfy_org,
+            },
+            node_id=cls.hidden.unique_id,
            prompt_text=prompt_text,
            seed=seed,
            model=model,
@@ -155,13 +204,17 @@ class MinimaxTextToVideoNode(IO.ComfyNode):


 class MinimaxImageToVideoNode(IO.ComfyNode):
+    """
+    Generates videos synchronously based on an image and prompt, and optional parameters using MiniMax's API.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="MinimaxImageToVideoNode",
            display_name="MiniMax Image to Video",
            category="api node/video/MiniMax",
-            description="Generates videos synchronously based on an image and prompt, and optional parameters.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.Image.Input(
                    "image",
@@ -208,7 +261,11 @@ class MinimaxImageToVideoNode(IO.ComfyNode):
        seed: int = 0,
    ) -> IO.NodeOutput:
        return await _generate_mm_video(
-            cls,
+            auth={
+                "auth_token": cls.hidden.auth_token_comfy_org,
+                "comfy_api_key": cls.hidden.api_key_comfy_org,
+            },
+            node_id=cls.hidden.unique_id,
            prompt_text=prompt_text,
            seed=seed,
            model=model,
@@ -219,13 +276,17 @@ class MinimaxImageToVideoNode(IO.ComfyNode):


 class MinimaxSubjectToVideoNode(IO.ComfyNode):
+    """
+    Generates videos synchronously based on an image and prompt, and optional parameters using MiniMax's API.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="MinimaxSubjectToVideoNode",
            display_name="MiniMax Subject to Video",
            category="api node/video/MiniMax",
-            description="Generates videos synchronously based on an image and prompt, and optional parameters.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.Image.Input(
                    "subject",
@@ -272,7 +333,11 @@ class MinimaxSubjectToVideoNode(IO.ComfyNode):
        seed: int = 0,
    ) -> IO.NodeOutput:
        return await _generate_mm_video(
-            cls,
+            auth={
+                "auth_token": cls.hidden.auth_token_comfy_org,
+                "comfy_api_key": cls.hidden.api_key_comfy_org,
+            },
+            node_id=cls.hidden.unique_id,
            prompt_text=prompt_text,
            seed=seed,
            model=model,
@@ -283,13 +348,15 @@ class MinimaxSubjectToVideoNode(IO.ComfyNode):


 class MinimaxHailuoVideoNode(IO.ComfyNode):
+    """Generates videos from prompt, with optional start frame using the new MiniMax Hailuo-02 model."""
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="MinimaxHailuoVideoNode",
            display_name="MiniMax Hailuo Video",
            category="api node/video/MiniMax",
-            description="Generates videos from prompt, with optional start frame using the new MiniMax Hailuo-02 model.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.String.Input(
                    "prompt_text",
@@ -353,6 +420,10 @@ class MinimaxHailuoVideoNode(IO.ComfyNode):
        resolution: str = "768P",
        model: str = "MiniMax-Hailuo-02",
    ) -> IO.NodeOutput:
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        if first_frame_image is None:
            validate_string(prompt_text, field_name="prompt_text")

@@ -364,13 +435,16 @@ class MinimaxHailuoVideoNode(IO.ComfyNode):
        # upload image, if passed in
        image_url = None
        if first_frame_image is not None:
-            image_url = (await upload_images_to_comfyapi(cls, first_frame_image, max_images=1))[0]
+            image_url = (await upload_images_to_comfyapi(first_frame_image, max_images=1, auth_kwargs=auth))[0]

-        response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/minimax/video_generation", method="POST"),
-            response_model=MinimaxVideoGenerationResponse,
-            data=MinimaxVideoGenerationRequest(
+        video_generate_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/minimax/video_generation",
+                method=HttpMethod.POST,
+                request_model=MinimaxVideoGenerationRequest,
+                response_model=MinimaxVideoGenerationResponse,
+            ),
+            request=MinimaxVideoGenerationRequest(
                model=MiniMaxModel(model),
                prompt=prompt_text,
                callback_url=None,
@@ -379,42 +453,67 @@ class MinimaxHailuoVideoNode(IO.ComfyNode):
                duration=duration,
                resolution=resolution,
            ),
+            auth_kwargs=auth,
        )
+        response = await video_generate_operation.execute()

        task_id = response.task_id
        if not task_id:
            raise Exception(f"MiniMax generation failed: {response.base_resp}")

        average_duration = 120 if resolution == "768P" else 240
-        task_result = await poll_op(
-            cls,
-            ApiEndpoint(path="/proxy/minimax/query/video_generation", query_params={"task_id": task_id}),
-            response_model=MinimaxTaskResultResponse,
+        video_generate_operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path="/proxy/minimax/query/video_generation",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=MinimaxTaskResultResponse,
+                query_params={"task_id": task_id},
+            ),
+            completed_statuses=["Success"],
+            failed_statuses=["Fail"],
            status_extractor=lambda x: x.status.value,
            estimated_duration=average_duration,
+            node_id=cls.hidden.unique_id,
+            auth_kwargs=auth,
        )
+        task_result = await video_generate_operation.execute()

        file_id = task_result.file_id
        if file_id is None:
            raise Exception("Request was not successful. Missing file ID.")
-        file_result = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/minimax/files/retrieve", query_params={"file_id": int(file_id)}),
-            response_model=MinimaxFileRetrieveResponse,
+        file_retrieve_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/minimax/files/retrieve",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=MinimaxFileRetrieveResponse,
+                query_params={"file_id": int(file_id)},
+            ),
+            request=EmptyRequest(),
+            auth_kwargs=auth,
        )
+        file_result = await file_retrieve_operation.execute()

        file_url = file_result.file.download_url
        if file_url is None:
-            raise Exception(f"No video was found in the response. Full response: {file_result.model_dump()}")
+            raise Exception(
+                f"No video was found in the response. Full response: {file_result.model_dump()}"
+            )
+        logging.info("Generated video URL: %s", file_url)
+        if cls.hidden.unique_id:
+            if hasattr(file_result.file, "backup_download_url"):
+                message = f"Result URL: {file_url}\nBackup URL: {file_result.file.backup_download_url}"
+            else:
+                message = f"Result URL: {file_url}"
+            PromptServer.instance.send_progress_text(message, cls.hidden.unique_id)

-        if file_result.file.backup_download_url:
-            try:
-                return IO.NodeOutput(await download_url_to_video_output(file_url, timeout=10, max_retries=2))
-            except Exception:  # if we have a second URL to retrieve the result, try again using that one
-                return IO.NodeOutput(
-                    await download_url_to_video_output(file_result.file.backup_download_url, max_retries=3)
-                )
-        return IO.NodeOutput(await download_url_to_video_output(file_url))
+        video_io = await download_url_to_bytesio(file_url)
+        if video_io is None:
+            error_msg = f"Failed to download video from {file_url}"
+            logging.error(error_msg)
+            raise Exception(error_msg)
+        return IO.NodeOutput(VideoFromFile(video_io))


 class MinimaxExtension(ComfyExtension):
--- a/comfy_api_nodes/nodes_openai.py
+++ b/comfy_api_nodes/nodes_openai.py
--- a/comfy_api_nodes/nodes_pika.py
+++ b/comfy_api_nodes/nodes_pika.py
@@ -7,23 +7,24 @@ from __future__ import annotations

 from io import BytesIO
 import logging
-from typing import Optional
+from typing import Optional, TypeVar

 import torch

 from typing_extensions import override
 from comfy_api.latest import ComfyExtension, IO
 from comfy_api.input_impl.video_types import VideoCodec, VideoContainer, VideoInput
-from comfy_api_nodes.apis import pika_api as pika_defs
-from comfy_api_nodes.util import (
-    validate_string,
-    download_url_to_video_output,
-    tensor_to_bytesio,
+from comfy_api_nodes.apis import pika_defs
+from comfy_api_nodes.apis.client import (
    ApiEndpoint,
-    sync_op,
-    poll_op,
+    EmptyRequest,
+    HttpMethod,
+    PollingOperation,
+    SynchronousOperation,
 )
+from comfy_api_nodes.util import validate_string, download_url_to_video_output, tensor_to_bytesio

+R = TypeVar("R")

 PATH_PIKADDITIONS = "/proxy/pika/generate/pikadditions"
 PATH_PIKASWAPS = "/proxy/pika/generate/pikaswaps"
@@ -39,18 +40,28 @@ PATH_VIDEO_GET = "/proxy/pika/videos"


 async def execute_task(
-    task_id: str,
-    cls: type[IO.ComfyNode],
+    initial_operation: SynchronousOperation[R, pika_defs.PikaGenerateResponse],
+    auth_kwargs: Optional[dict[str, str]] = None,
+    node_id: Optional[str] = None,
 ) -> IO.NodeOutput:
-    final_response: pika_defs.PikaVideoResponse = await poll_op(
-        cls,
-        ApiEndpoint(path=f"{PATH_VIDEO_GET}/{task_id}"),
-        response_model=pika_defs.PikaVideoResponse,
+    task_id = (await initial_operation.execute()).video_id
+    final_response: pika_defs.PikaVideoResponse = await PollingOperation(
+        poll_endpoint=ApiEndpoint(
+            path=f"{PATH_VIDEO_GET}/{task_id}",
+            method=HttpMethod.GET,
+            request_model=EmptyRequest,
+            response_model=pika_defs.PikaVideoResponse,
+        ),
+        completed_statuses=["finished"],
+        failed_statuses=["failed", "cancelled"],
        status_extractor=lambda response: (response.status.value if response.status else None),
        progress_extractor=lambda response: (response.progress if hasattr(response, "progress") else None),
+        auth_kwargs=auth_kwargs,
+        result_url_extractor=lambda response: (response.url if hasattr(response, "url") else None),
+        node_id=node_id,
        estimated_duration=60,
        max_poll_attempts=240,
-    )
+    ).execute()
    if not final_response.url:
        error_msg = f"Pika task {task_id} succeeded but no video data found in response:\n{final_response}"
        logging.error(error_msg)
@@ -113,15 +124,23 @@ class PikaImageToVideo(IO.ComfyNode):
            resolution=resolution,
            duration=duration,
        )
-        initial_operation = await sync_op(
-            cls,
-            ApiEndpoint(path=PATH_IMAGE_TO_VIDEO, method="POST"),
-            response_model=pika_defs.PikaGenerateResponse,
-            data=pika_request_data,
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        initial_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path=PATH_IMAGE_TO_VIDEO,
+                method=HttpMethod.POST,
+                request_model=pika_defs.PikaBodyGenerate22I2vGenerate22I2vPost,
+                response_model=pika_defs.PikaGenerateResponse,
+            ),
+            request=pika_request_data,
            files=pika_files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
-        return await execute_task(initial_operation.video_id, cls)
+        return await execute_task(initial_operation, auth_kwargs=auth, node_id=cls.hidden.unique_id)


 class PikaTextToVideoNode(IO.ComfyNode):
@@ -164,11 +183,18 @@ class PikaTextToVideoNode(IO.ComfyNode):
        duration: int,
        aspect_ratio: float,
    ) -> IO.NodeOutput:
-        initial_operation = await sync_op(
-            cls,
-            ApiEndpoint(path=PATH_TEXT_TO_VIDEO, method="POST"),
-            response_model=pika_defs.PikaGenerateResponse,
-            data=pika_defs.PikaBodyGenerate22T2vGenerate22T2vPost(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        initial_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path=PATH_TEXT_TO_VIDEO,
+                method=HttpMethod.POST,
+                request_model=pika_defs.PikaBodyGenerate22T2vGenerate22T2vPost,
+                response_model=pika_defs.PikaGenerateResponse,
+            ),
+            request=pika_defs.PikaBodyGenerate22T2vGenerate22T2vPost(
                promptText=prompt_text,
                negativePrompt=negative_prompt,
                seed=seed,
@@ -176,9 +202,10 @@ class PikaTextToVideoNode(IO.ComfyNode):
                duration=duration,
                aspectRatio=aspect_ratio,
            ),
+            auth_kwargs=auth,
            content_type="application/x-www-form-urlencoded",
        )
-        return await execute_task(initial_operation.video_id, cls)
+        return await execute_task(initial_operation, auth_kwargs=auth, node_id=cls.hidden.unique_id)


 class PikaScenes(IO.ComfyNode):
@@ -282,16 +309,24 @@ class PikaScenes(IO.ComfyNode):
            duration=duration,
            aspectRatio=aspect_ratio,
        )
-        initial_operation = await sync_op(
-            cls,
-            ApiEndpoint(path=PATH_PIKASCENES, method="POST"),
-            response_model=pika_defs.PikaGenerateResponse,
-            data=pika_request_data,
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        initial_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path=PATH_PIKASCENES,
+                method=HttpMethod.POST,
+                request_model=pika_defs.PikaBodyGenerate22C2vGenerate22PikascenesPost,
+                response_model=pika_defs.PikaGenerateResponse,
+            ),
+            request=pika_request_data,
            files=pika_files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )

-        return await execute_task(initial_operation.video_id, cls)
+        return await execute_task(initial_operation, auth_kwargs=auth, node_id=cls.hidden.unique_id)


 class PikAdditionsNode(IO.ComfyNode):
@@ -348,16 +383,24 @@ class PikAdditionsNode(IO.ComfyNode):
            negativePrompt=negative_prompt,
            seed=seed,
        )
-        initial_operation = await sync_op(
-            cls,
-            ApiEndpoint(path=PATH_PIKADDITIONS, method="POST"),
-            response_model=pika_defs.PikaGenerateResponse,
-            data=pika_request_data,
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        initial_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path=PATH_PIKADDITIONS,
+                method=HttpMethod.POST,
+                request_model=pika_defs.PikaBodyGeneratePikadditionsGeneratePikadditionsPost,
+                response_model=pika_defs.PikaGenerateResponse,
+            ),
+            request=pika_request_data,
            files=pika_files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )

-        return await execute_task(initial_operation.video_id, cls)
+        return await execute_task(initial_operation, auth_kwargs=auth, node_id=cls.hidden.unique_id)


 class PikaSwapsNode(IO.ComfyNode):
@@ -429,15 +472,23 @@ class PikaSwapsNode(IO.ComfyNode):
            seed=seed,
            modifyRegionRoi=region_to_modify if region_to_modify else None,
        )
-        initial_operation = await sync_op(
-            cls,
-            ApiEndpoint(path=PATH_PIKASWAPS, method="POST"),
-            response_model=pika_defs.PikaGenerateResponse,
-            data=pika_request_data,
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        initial_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path=PATH_PIKASWAPS,
+                method=HttpMethod.POST,
+                request_model=pika_defs.PikaBodyGeneratePikaswapsGeneratePikaswapsPost,
+                response_model=pika_defs.PikaGenerateResponse,
+            ),
+            request=pika_request_data,
            files=pika_files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
-        return await execute_task(initial_operation.video_id, cls)
+        return await execute_task(initial_operation, auth_kwargs=auth, node_id=cls.hidden.unique_id)


 class PikaffectsNode(IO.ComfyNode):
@@ -477,11 +528,18 @@ class PikaffectsNode(IO.ComfyNode):
        negative_prompt: str,
        seed: int,
    ) -> IO.NodeOutput:
-        initial_operation = await sync_op(
-            cls,
-            ApiEndpoint(path=PATH_PIKAFFECTS, method="POST"),
-            response_model=pika_defs.PikaGenerateResponse,
-            data=pika_defs.PikaBodyGeneratePikaffectsGeneratePikaffectsPost(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        initial_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path=PATH_PIKAFFECTS,
+                method=HttpMethod.POST,
+                request_model=pika_defs.PikaBodyGeneratePikaffectsGeneratePikaffectsPost,
+                response_model=pika_defs.PikaGenerateResponse,
+            ),
+            request=pika_defs.PikaBodyGeneratePikaffectsGeneratePikaffectsPost(
                pikaffect=pikaffect,
                promptText=prompt_text,
                negativePrompt=negative_prompt,
@@ -489,8 +547,9 @@ class PikaffectsNode(IO.ComfyNode):
            ),
            files={"image": ("image.png", tensor_to_bytesio(image), "image/png")},
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
-        return await execute_task(initial_operation.video_id, cls)
+        return await execute_task(initial_operation, auth_kwargs=auth, node_id=cls.hidden.unique_id)


 class PikaStartEndFrameNode(IO.ComfyNode):
@@ -533,11 +592,18 @@ class PikaStartEndFrameNode(IO.ComfyNode):
            ("keyFrames", ("image_start.png", tensor_to_bytesio(image_start), "image/png")),
            ("keyFrames", ("image_end.png", tensor_to_bytesio(image_end), "image/png")),
        ]
-        initial_operation = await sync_op(
-            cls,
-            ApiEndpoint(path=PATH_PIKAFRAMES, method="POST"),
-            response_model=pika_defs.PikaGenerateResponse,
-            data=pika_defs.PikaBodyGenerate22KeyframeGenerate22PikaframesPost(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        initial_operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path=PATH_PIKAFRAMES,
+                method=HttpMethod.POST,
+                request_model=pika_defs.PikaBodyGenerate22KeyframeGenerate22PikaframesPost,
+                response_model=pika_defs.PikaGenerateResponse,
+            ),
+            request=pika_defs.PikaBodyGenerate22KeyframeGenerate22PikaframesPost(
                promptText=prompt_text,
                negativePrompt=negative_prompt,
                seed=seed,
@@ -546,8 +612,9 @@ class PikaStartEndFrameNode(IO.ComfyNode):
            ),
            files=pika_files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
-        return await execute_task(initial_operation.video_id, cls)
+        return await execute_task(initial_operation, auth_kwargs=auth, node_id=cls.hidden.unique_id)


 class PikaApiNodesExtension(ComfyExtension):
--- a/comfy_api_nodes/nodes_pixverse.py
+++ b/comfy_api_nodes/nodes_pixverse.py
@@ -1,6 +1,7 @@
-import torch
+from inspect import cleandoc
+from typing import Optional
 from typing_extensions import override
-from comfy_api.latest import IO, ComfyExtension
+from io import BytesIO
 from comfy_api_nodes.apis.pixverse_api import (
    PixverseTextVideoRequest,
    PixverseImageVideoRequest,
@@ -16,30 +17,53 @@ from comfy_api_nodes.apis.pixverse_api import (
    PixverseIO,
    pixverse_templates,
 )
-from comfy_api_nodes.util import (
+from comfy_api_nodes.apis.client import (
    ApiEndpoint,
-    download_url_to_video_output,
-    poll_op,
-    sync_op,
-    tensor_to_bytesio,
-    validate_string,
+    HttpMethod,
+    SynchronousOperation,
+    PollingOperation,
+    EmptyRequest,
 )
+from comfy_api_nodes.util import validate_string, tensor_to_bytesio
+from comfy_api.input_impl import VideoFromFile
+from comfy_api.latest import ComfyExtension, IO
+
+import torch
+import aiohttp
+

 AVERAGE_DURATION_T2V = 32
 AVERAGE_DURATION_I2V = 30
 AVERAGE_DURATION_T2T = 52


-async def upload_image_to_pixverse(cls: type[IO.ComfyNode], image: torch.Tensor):
-    response_upload = await sync_op(
-        cls,
-        ApiEndpoint(path="/proxy/pixverse/image/upload", method="POST"),
-        response_model=PixverseImageUploadResponse,
+def get_video_url_from_response(
+    response: PixverseGenerationStatusResponse,
+) -> Optional[str]:
+    if response.Resp is None or response.Resp.url is None:
+        return None
+    return str(response.Resp.url)
+
+
+async def upload_image_to_pixverse(image: torch.Tensor, auth_kwargs=None):
+    # first, upload image to Pixverse and get image id to use in actual generation call
+    operation = SynchronousOperation(
+        endpoint=ApiEndpoint(
+            path="/proxy/pixverse/image/upload",
+            method=HttpMethod.POST,
+            request_model=EmptyRequest,
+            response_model=PixverseImageUploadResponse,
+        ),
+        request=EmptyRequest(),
        files={"image": tensor_to_bytesio(image)},
        content_type="multipart/form-data",
+        auth_kwargs=auth_kwargs,
    )
+    response_upload: PixverseImageUploadResponse = await operation.execute()
+
    if response_upload.Resp is None:
        raise Exception(f"PixVerse image upload request failed: '{response_upload.ErrMsg}'")
+
    return response_upload.Resp.img_id


@@ -69,13 +93,17 @@ class PixverseTemplateNode(IO.ComfyNode):


 class PixverseTextToVideoNode(IO.ComfyNode):
+    """
+    Generates videos based on prompt and output_size.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="PixverseTextToVideoNode",
            display_name="PixVerse Text to Video",
            category="api node/video/PixVerse",
-            description="Generates videos based on prompt and output_size.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.String.Input(
                    "prompt",
@@ -142,7 +170,7 @@ class PixverseTextToVideoNode(IO.ComfyNode):
        negative_prompt: str = None,
        pixverse_template: int = None,
    ) -> IO.NodeOutput:
-        validate_string(prompt, strip_whitespace=False, min_length=1)
+        validate_string(prompt, strip_whitespace=False)
        # 1080p is limited to 5 seconds duration
        # only normal motion_mode supported for 1080p or for non-5 second duration
        if quality == PixverseQuality.res_1080p:
@@ -151,11 +179,18 @@ class PixverseTextToVideoNode(IO.ComfyNode):
        elif duration_seconds != PixverseDuration.dur_5:
            motion_mode = PixverseMotionMode.normal

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/pixverse/video/text/generate", method="POST"),
-            response_model=PixverseVideoResponse,
-            data=PixverseTextVideoRequest(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/pixverse/video/text/generate",
+                method=HttpMethod.POST,
+                request_model=PixverseTextVideoRequest,
+                response_model=PixverseVideoResponse,
+            ),
+            request=PixverseTextVideoRequest(
                prompt=prompt,
                aspect_ratio=aspect_ratio,
                quality=quality,
@@ -165,14 +200,20 @@ class PixverseTextToVideoNode(IO.ComfyNode):
                template_id=pixverse_template,
                seed=seed,
            ),
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()
+
        if response_api.Resp is None:
            raise Exception(f"PixVerse request failed: '{response_api.ErrMsg}'")

-        response_poll = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/pixverse/video/result/{response_api.Resp.video_id}"),
-            response_model=PixverseGenerationStatusResponse,
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/pixverse/video/result/{response_api.Resp.video_id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=PixverseGenerationStatusResponse,
+            ),
            completed_statuses=[PixverseStatus.successful],
            failed_statuses=[
                PixverseStatus.contents_moderation,
@@ -180,19 +221,30 @@ class PixverseTextToVideoNode(IO.ComfyNode):
                PixverseStatus.deleted,
            ],
            status_extractor=lambda x: x.Resp.status,
+            auth_kwargs=auth,
+            node_id=cls.hidden.unique_id,
+            result_url_extractor=get_video_url_from_response,
            estimated_duration=AVERAGE_DURATION_T2V,
        )
-        return IO.NodeOutput(await download_url_to_video_output(response_poll.Resp.url))
+        response_poll = await operation.execute()
+
+        async with aiohttp.ClientSession() as session:
+            async with session.get(response_poll.Resp.url) as vid_response:
+                return IO.NodeOutput(VideoFromFile(BytesIO(await vid_response.content.read())))


 class PixverseImageToVideoNode(IO.ComfyNode):
+    """
+    Generates videos based on prompt and output_size.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="PixverseImageToVideoNode",
            display_name="PixVerse Image to Video",
            category="api node/video/PixVerse",
-            description="Generates videos based on prompt and output_size.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.Image.Input("image"),
                IO.String.Input(
@@ -257,7 +309,11 @@ class PixverseImageToVideoNode(IO.ComfyNode):
        pixverse_template: int = None,
    ) -> IO.NodeOutput:
        validate_string(prompt, strip_whitespace=False)
-        img_id = await upload_image_to_pixverse(cls, image)
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        img_id = await upload_image_to_pixverse(image, auth_kwargs=auth)

        # 1080p is limited to 5 seconds duration
        # only normal motion_mode supported for 1080p or for non-5 second duration
@@ -267,11 +323,14 @@ class PixverseImageToVideoNode(IO.ComfyNode):
        elif duration_seconds != PixverseDuration.dur_5:
            motion_mode = PixverseMotionMode.normal

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/pixverse/video/img/generate", method="POST"),
-            response_model=PixverseVideoResponse,
-            data=PixverseImageVideoRequest(
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/pixverse/video/img/generate",
+                method=HttpMethod.POST,
+                request_model=PixverseImageVideoRequest,
+                response_model=PixverseVideoResponse,
+            ),
+            request=PixverseImageVideoRequest(
                img_id=img_id,
                prompt=prompt,
                quality=quality,
@@ -281,15 +340,20 @@ class PixverseImageToVideoNode(IO.ComfyNode):
                template_id=pixverse_template,
                seed=seed,
            ),
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()

        if response_api.Resp is None:
            raise Exception(f"PixVerse request failed: '{response_api.ErrMsg}'")

-        response_poll = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/pixverse/video/result/{response_api.Resp.video_id}"),
-            response_model=PixverseGenerationStatusResponse,
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/pixverse/video/result/{response_api.Resp.video_id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=PixverseGenerationStatusResponse,
+            ),
            completed_statuses=[PixverseStatus.successful],
            failed_statuses=[
                PixverseStatus.contents_moderation,
@@ -297,19 +361,30 @@ class PixverseImageToVideoNode(IO.ComfyNode):
                PixverseStatus.deleted,
            ],
            status_extractor=lambda x: x.Resp.status,
+            auth_kwargs=auth,
+            node_id=cls.hidden.unique_id,
+            result_url_extractor=get_video_url_from_response,
            estimated_duration=AVERAGE_DURATION_I2V,
        )
-        return IO.NodeOutput(await download_url_to_video_output(response_poll.Resp.url))
+        response_poll = await operation.execute()
+
+        async with aiohttp.ClientSession() as session:
+            async with session.get(response_poll.Resp.url) as vid_response:
+                return IO.NodeOutput(VideoFromFile(BytesIO(await vid_response.content.read())))


 class PixverseTransitionVideoNode(IO.ComfyNode):
+    """
+    Generates videos based on prompt and output_size.
+    """
+
    @classmethod
    def define_schema(cls) -> IO.Schema:
        return IO.Schema(
            node_id="PixverseTransitionVideoNode",
            display_name="PixVerse Transition Video",
            category="api node/video/PixVerse",
-            description="Generates videos based on prompt and output_size.",
+            description=cleandoc(cls.__doc__ or ""),
            inputs=[
                IO.Image.Input("first_frame"),
                IO.Image.Input("last_frame"),
@@ -370,8 +445,12 @@ class PixverseTransitionVideoNode(IO.ComfyNode):
        negative_prompt: str = None,
    ) -> IO.NodeOutput:
        validate_string(prompt, strip_whitespace=False)
-        first_frame_id = await upload_image_to_pixverse(cls, first_frame)
-        last_frame_id = await upload_image_to_pixverse(cls, last_frame)
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+        first_frame_id = await upload_image_to_pixverse(first_frame, auth_kwargs=auth)
+        last_frame_id = await upload_image_to_pixverse(last_frame, auth_kwargs=auth)

        # 1080p is limited to 5 seconds duration
        # only normal motion_mode supported for 1080p or for non-5 second duration
@@ -381,11 +460,14 @@ class PixverseTransitionVideoNode(IO.ComfyNode):
        elif duration_seconds != PixverseDuration.dur_5:
            motion_mode = PixverseMotionMode.normal

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/pixverse/video/transition/generate", method="POST"),
-            response_model=PixverseVideoResponse,
-            data=PixverseTransitionVideoRequest(
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/pixverse/video/transition/generate",
+                method=HttpMethod.POST,
+                request_model=PixverseTransitionVideoRequest,
+                response_model=PixverseVideoResponse,
+            ),
+            request=PixverseTransitionVideoRequest(
                first_frame_img=first_frame_id,
                last_frame_img=last_frame_id,
                prompt=prompt,
@@ -395,15 +477,20 @@ class PixverseTransitionVideoNode(IO.ComfyNode):
                negative_prompt=negative_prompt if negative_prompt else None,
                seed=seed,
            ),
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()

        if response_api.Resp is None:
            raise Exception(f"PixVerse request failed: '{response_api.ErrMsg}'")

-        response_poll = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/pixverse/video/result/{response_api.Resp.video_id}"),
-            response_model=PixverseGenerationStatusResponse,
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/pixverse/video/result/{response_api.Resp.video_id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=PixverseGenerationStatusResponse,
+            ),
            completed_statuses=[PixverseStatus.successful],
            failed_statuses=[
                PixverseStatus.contents_moderation,
@@ -411,9 +498,16 @@ class PixverseTransitionVideoNode(IO.ComfyNode):
                PixverseStatus.deleted,
            ],
            status_extractor=lambda x: x.Resp.status,
+            auth_kwargs=auth,
+            node_id=cls.hidden.unique_id,
+            result_url_extractor=get_video_url_from_response,
            estimated_duration=AVERAGE_DURATION_T2V,
        )
-        return IO.NodeOutput(await download_url_to_video_output(response_poll.Resp.url))
+        response_poll = await operation.execute()
+
+        async with aiohttp.ClientSession() as session:
+            async with session.get(response_poll.Resp.url) as vid_response:
+                return IO.NodeOutput(VideoFromFile(BytesIO(await vid_response.content.read())))


 class PixVerseExtension(ComfyExtension):
--- a/comfy_api_nodes/nodes_recraft.py
+++ b/comfy_api_nodes/nodes_recraft.py
@@ -8,6 +8,9 @@ from typing_extensions import override

 from comfy.utils import ProgressBar
 from comfy_api.latest import IO, ComfyExtension
+from comfy_api_nodes.apinode_utils import (
+    resize_mask_to_image,
+)
 from comfy_api_nodes.apis.recraft_api import (
    RecraftColor,
    RecraftColorChain,
@@ -25,7 +28,6 @@ from comfy_api_nodes.util import (
    ApiEndpoint,
    bytesio_to_image_tensor,
    download_url_as_bytesio,
-    resize_mask_to_image,
    sync_op,
    tensor_to_bytesio,
    validate_string,
--- a/comfy_api_nodes/nodes_rodin.py
+++ b/comfy_api_nodes/nodes_rodin.py
@@ -5,9 +5,12 @@ Rodin API docs: https://developer.hyper3d.ai/

 """

+from __future__ import annotations
 from inspect import cleandoc
 import folder_paths as comfy_paths
+import aiohttp
 import os
+import asyncio
 import logging
 import math
 from typing import Optional
@@ -23,11 +26,11 @@ from comfy_api_nodes.apis.rodin_api import (
    Rodin3DDownloadResponse,
    JobStatus,
 )
-from comfy_api_nodes.util import (
-    sync_op,
-    poll_op,
+from comfy_api_nodes.apis.client import (
    ApiEndpoint,
-    download_url_to_bytesio,
+    HttpMethod,
+    SynchronousOperation,
+    PollingOperation,
 )
 from comfy_api.latest import ComfyExtension, IO

@@ -118,31 +121,35 @@ def tensor_to_filelike(tensor, max_pixels: int = 2048*2048):


 async def create_generate_task(
-    cls: type[IO.ComfyNode],
    images=None,
    seed=1,
    material="PBR",
    quality_override=18000,
    tier="Regular",
    mesh_mode="Quad",
-    ta_pose: bool = False,
+    TAPose = False,
+    auth_kwargs: Optional[dict[str, str]] = None,
 ):
    if images is None:
        raise Exception("Rodin 3D generate requires at least 1 image.")
    if len(images) > 5:
        raise Exception("Rodin 3D generate requires up to 5 image.")

-    response = await sync_op(
-        cls,
-        ApiEndpoint(path="/proxy/rodin/api/v2/rodin", method="POST"),
-        response_model=Rodin3DGenerateResponse,
-        data=Rodin3DGenerateRequest(
+    path = "/proxy/rodin/api/v2/rodin"
+    operation = SynchronousOperation(
+        endpoint=ApiEndpoint(
+            path=path,
+            method=HttpMethod.POST,
+            request_model=Rodin3DGenerateRequest,
+            response_model=Rodin3DGenerateResponse,
+        ),
+        request=Rodin3DGenerateRequest(
            seed=seed,
            tier=tier,
            material=material,
            quality_override=quality_override,
            mesh_mode=mesh_mode,
-            TAPose=ta_pose,
+            TAPose=TAPose,
        ),
        files=[
            (
@@ -152,8 +159,11 @@ async def create_generate_task(
            for image in images if image is not None
        ],
        content_type="multipart/form-data",
+        auth_kwargs=auth_kwargs,
    )

+    response = await operation.execute()
+
    if hasattr(response, "error"):
        error_message = f"Rodin3D Create 3D generate Task Failed. Message: {response.message}, error: {response.error}"
        logging.error(error_message)
@@ -177,46 +187,75 @@ def check_rodin_status(response: Rodin3DCheckStatusResponse) -> str:
        return "DONE"
    return "Generating"

-def extract_progress(response: Rodin3DCheckStatusResponse) -> Optional[int]:
-    if not response.jobs:
-        return None
-    completed_count = sum(1 for job in response.jobs if job.status == JobStatus.Done)
-    return int((completed_count / len(response.jobs)) * 100)

-
-async def poll_for_task_status(subscription_key: str, cls: type[IO.ComfyNode]) -> Rodin3DCheckStatusResponse:
-    logging.info("[ Rodin3D API - CheckStatus ] Generate Start!")
-    return await poll_op(
-        cls,
-        ApiEndpoint(path="/proxy/rodin/api/v2/status", method="POST"),
-        response_model=Rodin3DCheckStatusResponse,
-        data=Rodin3DCheckStatusRequest(subscription_key=subscription_key),
+async def poll_for_task_status(
+    subscription_key, auth_kwargs: Optional[dict[str, str]] = None,
+) -> Rodin3DCheckStatusResponse:
+    poll_operation = PollingOperation(
+        poll_endpoint=ApiEndpoint(
+            path="/proxy/rodin/api/v2/status",
+            method=HttpMethod.POST,
+            request_model=Rodin3DCheckStatusRequest,
+            response_model=Rodin3DCheckStatusResponse,
+        ),
+        request=Rodin3DCheckStatusRequest(subscription_key=subscription_key),
+        completed_statuses=["DONE"],
+        failed_statuses=["FAILED"],
        status_extractor=check_rodin_status,
-        progress_extractor=extract_progress,
+        poll_interval=3.0,
+        auth_kwargs=auth_kwargs,
    )
+    logging.info("[ Rodin3D API - CheckStatus ] Generate Start!")
+    return await poll_operation.execute()


-async def get_rodin_download_list(uuid: str, cls: type[IO.ComfyNode]) -> Rodin3DDownloadResponse:
+async def get_rodin_download_list(uuid, auth_kwargs: Optional[dict[str, str]] = None) -> Rodin3DDownloadResponse:
    logging.info("[ Rodin3D API - Downloading ] Generate Successfully!")
-    return await sync_op(
-        cls,
-        ApiEndpoint(path="/proxy/rodin/api/v2/download", method="POST"),
-        response_model=Rodin3DDownloadResponse,
-        data=Rodin3DDownloadRequest(task_uuid=uuid),
-        monitor_progress=False,
+    operation = SynchronousOperation(
+        endpoint=ApiEndpoint(
+            path="/proxy/rodin/api/v2/download",
+            method=HttpMethod.POST,
+            request_model=Rodin3DDownloadRequest,
+            response_model=Rodin3DDownloadResponse,
+        ),
+        request=Rodin3DDownloadRequest(task_uuid=uuid),
+        auth_kwargs=auth_kwargs,
    )
+    return await operation.execute()


-async def download_files(url_list, task_uuid: str):
-    result_folder_name = f"Rodin3D_{task_uuid}"
-    save_path = os.path.join(comfy_paths.get_output_directory(), result_folder_name)
+async def download_files(url_list, task_uuid):
+    save_path = os.path.join(comfy_paths.get_output_directory(), f"Rodin3D_{task_uuid}")
    os.makedirs(save_path, exist_ok=True)
    model_file_path = None
-    for i in url_list.list:
-        file_path = os.path.join(save_path, i.name)
-        if file_path.endswith(".glb"):
-            model_file_path = os.path.join(result_folder_name, i.name)
-        await download_url_to_bytesio(i.url, file_path)
+    async with aiohttp.ClientSession() as session:
+        for i in url_list.list:
+            url = i.url
+            file_name = i.name
+            file_path = os.path.join(save_path, file_name)
+            if file_path.endswith(".glb"):
+                model_file_path = file_path
+            logging.info("[ Rodin3D API - download_files ] Downloading file: %s", file_path)
+            max_retries = 5
+            for attempt in range(max_retries):
+                try:
+                    async with session.get(url) as resp:
+                        resp.raise_for_status()
+                        with open(file_path, "wb") as f:
+                            async for chunk in resp.content.iter_chunked(32 * 1024):
+                                f.write(chunk)
+                    break
+                except Exception as e:
+                    logging.info("[ Rodin3D API - download_files ] Error downloading %s:%s", file_path, str(e))
+                    if attempt < max_retries - 1:
+                        logging.info("Retrying...")
+                        await asyncio.sleep(2)
+                    else:
+                        logging.info(
+                            "[ Rodin3D API - download_files ] Failed to download %s after %s attempts.",
+                            file_path,
+                            max_retries,
+                        )
    return model_file_path


@@ -238,7 +277,6 @@ class Rodin3D_Regular(IO.ComfyNode):
            hidden=[
                IO.Hidden.auth_token_comfy_org,
                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
            ],
            is_api_node=True,
        )
@@ -257,17 +295,21 @@ class Rodin3D_Regular(IO.ComfyNode):
        for i in range(num_images):
            m_images.append(Images[i])
        mesh_mode, quality_override = get_quality_mode(Polygon_count)
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        task_uuid, subscription_key = await create_generate_task(
-            cls,
            images=m_images,
            seed=Seed,
            material=Material_Type,
            quality_override=quality_override,
            tier=tier,
            mesh_mode=mesh_mode,
+            auth_kwargs=auth,
        )
-        await poll_for_task_status(subscription_key, cls)
-        download_list = await get_rodin_download_list(task_uuid, cls)
+        await poll_for_task_status(subscription_key, auth_kwargs=auth)
+        download_list = await get_rodin_download_list(task_uuid, auth_kwargs=auth)
        model = await download_files(download_list, task_uuid)

        return IO.NodeOutput(model)
@@ -291,7 +333,6 @@ class Rodin3D_Detail(IO.ComfyNode):
            hidden=[
                IO.Hidden.auth_token_comfy_org,
                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
            ],
            is_api_node=True,
        )
@@ -310,17 +351,21 @@ class Rodin3D_Detail(IO.ComfyNode):
        for i in range(num_images):
            m_images.append(Images[i])
        mesh_mode, quality_override = get_quality_mode(Polygon_count)
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        task_uuid, subscription_key = await create_generate_task(
-            cls,
            images=m_images,
            seed=Seed,
            material=Material_Type,
            quality_override=quality_override,
            tier=tier,
            mesh_mode=mesh_mode,
+            auth_kwargs=auth,
        )
-        await poll_for_task_status(subscription_key, cls)
-        download_list = await get_rodin_download_list(task_uuid, cls)
+        await poll_for_task_status(subscription_key, auth_kwargs=auth)
+        download_list = await get_rodin_download_list(task_uuid, auth_kwargs=auth)
        model = await download_files(download_list, task_uuid)

        return IO.NodeOutput(model)
@@ -344,7 +389,6 @@ class Rodin3D_Smooth(IO.ComfyNode):
            hidden=[
                IO.Hidden.auth_token_comfy_org,
                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
            ],
            is_api_node=True,
        )
@@ -357,22 +401,27 @@ class Rodin3D_Smooth(IO.ComfyNode):
        Material_Type,
        Polygon_count,
    ) -> IO.NodeOutput:
+        tier = "Smooth"
        num_images = Images.shape[0]
        m_images = []
        for i in range(num_images):
            m_images.append(Images[i])
        mesh_mode, quality_override = get_quality_mode(Polygon_count)
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        task_uuid, subscription_key = await create_generate_task(
-            cls,
            images=m_images,
            seed=Seed,
            material=Material_Type,
            quality_override=quality_override,
-            tier="Smooth",
+            tier=tier,
            mesh_mode=mesh_mode,
+            auth_kwargs=auth,
        )
-        await poll_for_task_status(subscription_key, cls)
-        download_list = await get_rodin_download_list(task_uuid, cls)
+        await poll_for_task_status(subscription_key, auth_kwargs=auth)
+        download_list = await get_rodin_download_list(task_uuid, auth_kwargs=auth)
        model = await download_files(download_list, task_uuid)

        return IO.NodeOutput(model)
@@ -403,7 +452,6 @@ class Rodin3D_Sketch(IO.ComfyNode):
            hidden=[
                IO.Hidden.auth_token_comfy_org,
                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
            ],
            is_api_node=True,
        )
@@ -414,21 +462,29 @@ class Rodin3D_Sketch(IO.ComfyNode):
        Images,
        Seed,
    ) -> IO.NodeOutput:
+        tier = "Sketch"
        num_images = Images.shape[0]
        m_images = []
        for i in range(num_images):
            m_images.append(Images[i])
+        material_type = "PBR"
+        quality_override = 18000
+        mesh_mode = "Quad"
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        task_uuid, subscription_key = await create_generate_task(
-            cls,
            images=m_images,
            seed=Seed,
-            material="PBR",
-            quality_override=18000,
-            tier="Sketch",
-            mesh_mode="Quad",
+            material=material_type,
+            quality_override=quality_override,
+            tier=tier,
+            mesh_mode=mesh_mode,
+            auth_kwargs=auth,
        )
-        await poll_for_task_status(subscription_key, cls)
-        download_list = await get_rodin_download_list(task_uuid, cls)
+        await poll_for_task_status(subscription_key, auth_kwargs=auth)
+        download_list = await get_rodin_download_list(task_uuid, auth_kwargs=auth)
        model = await download_files(download_list, task_uuid)

        return IO.NodeOutput(model)
@@ -467,7 +523,6 @@ class Rodin3D_Gen2(IO.ComfyNode):
            hidden=[
                IO.Hidden.auth_token_comfy_org,
                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
            ],
            is_api_node=True,
        )
@@ -487,18 +542,22 @@ class Rodin3D_Gen2(IO.ComfyNode):
        for i in range(num_images):
            m_images.append(Images[i])
        mesh_mode, quality_override = get_quality_mode(Polygon_count)
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
        task_uuid, subscription_key = await create_generate_task(
-            cls,
            images=m_images,
            seed=Seed,
            material=Material_Type,
            quality_override=quality_override,
            tier=tier,
            mesh_mode=mesh_mode,
-            ta_pose=TAPose,
+            TAPose=TAPose,
+            auth_kwargs=auth,
        )
-        await poll_for_task_status(subscription_key, cls)
-        download_list = await get_rodin_download_list(task_uuid, cls)
+        await poll_for_task_status(subscription_key, auth_kwargs=auth)
+        download_list = await get_rodin_download_list(task_uuid, auth_kwargs=auth)
        model = await download_files(download_list, task_uuid)

        return IO.NodeOutput(model)
--- a/comfy_api_nodes/nodes_runway.py
+++ b/comfy_api_nodes/nodes_runway.py
@@ -200,7 +200,7 @@ class RunwayImageToVideoNodeGen3a(IO.ComfyNode):
    ) -> IO.NodeOutput:
        validate_string(prompt, min_length=1)
        validate_image_dimensions(start_frame, max_width=7999, max_height=7999)
-        validate_image_aspect_ratio(start_frame, (1, 2), (2, 1))
+        validate_image_aspect_ratio(start_frame, min_aspect_ratio=0.5, max_aspect_ratio=2.0)

        download_urls = await upload_images_to_comfyapi(
            cls,
@@ -290,7 +290,7 @@ class RunwayImageToVideoNodeGen4(IO.ComfyNode):
    ) -> IO.NodeOutput:
        validate_string(prompt, min_length=1)
        validate_image_dimensions(start_frame, max_width=7999, max_height=7999)
-        validate_image_aspect_ratio(start_frame, (1, 2), (2, 1))
+        validate_image_aspect_ratio(start_frame, min_aspect_ratio=0.5, max_aspect_ratio=2.0)

        download_urls = await upload_images_to_comfyapi(
            cls,
@@ -390,8 +390,8 @@ class RunwayFirstLastFrameNode(IO.ComfyNode):
        validate_string(prompt, min_length=1)
        validate_image_dimensions(start_frame, max_width=7999, max_height=7999)
        validate_image_dimensions(end_frame, max_width=7999, max_height=7999)
-        validate_image_aspect_ratio(start_frame, (1, 2), (2, 1))
-        validate_image_aspect_ratio(end_frame, (1, 2), (2, 1))
+        validate_image_aspect_ratio(start_frame, min_aspect_ratio=0.5, max_aspect_ratio=2.0)
+        validate_image_aspect_ratio(end_frame, min_aspect_ratio=0.5, max_aspect_ratio=2.0)

        stacked_input_images = image_tensor_pair_to_batch(start_frame, end_frame)
        download_urls = await upload_images_to_comfyapi(
@@ -475,7 +475,7 @@ class RunwayTextToImageNode(IO.ComfyNode):
        reference_images = None
        if reference_image is not None:
            validate_image_dimensions(reference_image, max_width=7999, max_height=7999)
-            validate_image_aspect_ratio(reference_image, (1, 2), (2, 1))
+            validate_image_aspect_ratio(reference_image, min_aspect_ratio=0.5, max_aspect_ratio=2.0)
            download_urls = await upload_images_to_comfyapi(
                cls,
                reference_image,
--- a/comfy_api_nodes/nodes_stability.py
+++ b/comfy_api_nodes/nodes_stability.py
@@ -20,6 +20,13 @@ from comfy_api_nodes.apis.stability_api import (
    StabilityAudioInpaintRequest,
    StabilityAudioResponse,
 )
+from comfy_api_nodes.apis.client import (
+    ApiEndpoint,
+    HttpMethod,
+    SynchronousOperation,
+    PollingOperation,
+    EmptyRequest,
+)
 from comfy_api_nodes.util import (
    validate_audio_duration,
    validate_string,
@@ -27,9 +34,6 @@ from comfy_api_nodes.util import (
    bytesio_to_image_tensor,
    tensor_to_bytesio,
    audio_bytes_to_audio_input,
-    sync_op,
-    poll_op,
-    ApiEndpoint,
 )

 import torch
@@ -157,11 +161,19 @@ class StabilityStableImageUltraNode(IO.ComfyNode):
            "image": image_binary
        }

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/stability/v2beta/stable-image/generate/ultra", method="POST"),
-            response_model=StabilityStableUltraResponse,
-            data=StabilityStableUltraRequest(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/stable-image/generate/ultra",
+                method=HttpMethod.POST,
+                request_model=StabilityStableUltraRequest,
+                response_model=StabilityStableUltraResponse,
+            ),
+            request=StabilityStableUltraRequest(
                prompt=prompt,
                negative_prompt=negative_prompt,
                aspect_ratio=aspect_ratio,
@@ -171,7 +183,9 @@ class StabilityStableImageUltraNode(IO.ComfyNode):
            ),
            files=files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()

        if response_api.finish_reason != "SUCCESS":
            raise Exception(f"Stable Image Ultra generation failed: {response_api.finish_reason}.")
@@ -299,11 +313,19 @@ class StabilityStableImageSD_3_5Node(IO.ComfyNode):
            "image": image_binary
        }

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/stability/v2beta/stable-image/generate/sd3", method="POST"),
-            response_model=StabilityStableUltraResponse,
-            data=StabilityStable3_5Request(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/stable-image/generate/sd3",
+                method=HttpMethod.POST,
+                request_model=StabilityStable3_5Request,
+                response_model=StabilityStableUltraResponse,
+            ),
+            request=StabilityStable3_5Request(
                prompt=prompt,
                negative_prompt=negative_prompt,
                aspect_ratio=aspect_ratio,
@@ -316,7 +338,9 @@ class StabilityStableImageSD_3_5Node(IO.ComfyNode):
            ),
            files=files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()

        if response_api.finish_reason != "SUCCESS":
            raise Exception(f"Stable Diffusion 3.5 Image generation failed: {response_api.finish_reason}.")
@@ -403,11 +427,19 @@ class StabilityUpscaleConservativeNode(IO.ComfyNode):
            "image": image_binary
        }

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/stability/v2beta/stable-image/upscale/conservative", method="POST"),
-            response_model=StabilityStableUltraResponse,
-            data=StabilityUpscaleConservativeRequest(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/stable-image/upscale/conservative",
+                method=HttpMethod.POST,
+                request_model=StabilityUpscaleConservativeRequest,
+                response_model=StabilityStableUltraResponse,
+            ),
+            request=StabilityUpscaleConservativeRequest(
                prompt=prompt,
                negative_prompt=negative_prompt,
                creativity=round(creativity,2),
@@ -415,7 +447,9 @@ class StabilityUpscaleConservativeNode(IO.ComfyNode):
            ),
            files=files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()

        if response_api.finish_reason != "SUCCESS":
            raise Exception(f"Stability Upscale Conservative generation failed: {response_api.finish_reason}.")
@@ -510,11 +544,19 @@ class StabilityUpscaleCreativeNode(IO.ComfyNode):
            "image": image_binary
        }

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/stability/v2beta/stable-image/upscale/creative", method="POST"),
-            response_model=StabilityAsyncResponse,
-            data=StabilityUpscaleCreativeRequest(
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/stable-image/upscale/creative",
+                method=HttpMethod.POST,
+                request_model=StabilityUpscaleCreativeRequest,
+                response_model=StabilityAsyncResponse,
+            ),
+            request=StabilityUpscaleCreativeRequest(
                prompt=prompt,
                negative_prompt=negative_prompt,
                creativity=round(creativity,2),
@@ -523,15 +565,25 @@ class StabilityUpscaleCreativeNode(IO.ComfyNode):
            ),
            files=files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()

-        response_poll = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/stability/v2beta/results/{response_api.id}"),
-            response_model=StabilityResultsGetResponse,
+        operation = PollingOperation(
+            poll_endpoint=ApiEndpoint(
+                path=f"/proxy/stability/v2beta/results/{response_api.id}",
+                method=HttpMethod.GET,
+                request_model=EmptyRequest,
+                response_model=StabilityResultsGetResponse,
+            ),
            poll_interval=3,
+            completed_statuses=[StabilityPollStatus.finished],
+            failed_statuses=[StabilityPollStatus.failed],
            status_extractor=lambda x: get_async_dummy_status(x),
+            auth_kwargs=auth,
+            node_id=cls.hidden.unique_id,
        )
+        response_poll: StabilityResultsGetResponse = await operation.execute()

        if response_poll.finish_reason != "SUCCESS":
            raise Exception(f"Stability Upscale Creative generation failed: {response_poll.finish_reason}.")
@@ -576,13 +628,24 @@ class StabilityUpscaleFastNode(IO.ComfyNode):
            "image": image_binary
        }

-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/stability/v2beta/stable-image/upscale/fast", method="POST"),
-            response_model=StabilityStableUltraResponse,
+        auth = {
+            "auth_token": cls.hidden.auth_token_comfy_org,
+            "comfy_api_key": cls.hidden.api_key_comfy_org,
+        }
+
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/stable-image/upscale/fast",
+                method=HttpMethod.POST,
+                request_model=EmptyRequest,
+                response_model=StabilityStableUltraResponse,
+            ),
+            request=EmptyRequest(),
            files=files,
            content_type="multipart/form-data",
+            auth_kwargs=auth,
        )
+        response_api = await operation.execute()

        if response_api.finish_reason != "SUCCESS":
            raise Exception(f"Stability Upscale Fast failed: {response_api.finish_reason}.")
@@ -654,13 +717,21 @@ class StabilityTextToAudio(IO.ComfyNode):
    async def execute(cls, model: str, prompt: str, duration: int, seed: int, steps: int) -> IO.NodeOutput:
        validate_string(prompt, max_length=10000)
        payload = StabilityTextToAudioRequest(prompt=prompt, model=model, duration=duration, seed=seed, steps=steps)
-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/stability/v2beta/audio/stable-audio-2/text-to-audio", method="POST"),
-            response_model=StabilityAudioResponse,
-            data=payload,
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/audio/stable-audio-2/text-to-audio",
+                method=HttpMethod.POST,
+                request_model=StabilityTextToAudioRequest,
+                response_model=StabilityAudioResponse,
+            ),
+            request=payload,
            content_type="multipart/form-data",
+            auth_kwargs= {
+                "auth_token": cls.hidden.auth_token_comfy_org,
+                "comfy_api_key": cls.hidden.api_key_comfy_org,
+            },
        )
+        response_api = await operation.execute()
        if not response_api.audio:
            raise ValueError("No audio file was received in response.")
        return IO.NodeOutput(audio_bytes_to_audio_input(base64.b64decode(response_api.audio)))
@@ -743,14 +814,22 @@ class StabilityAudioToAudio(IO.ComfyNode):
        payload = StabilityAudioToAudioRequest(
            prompt=prompt, model=model, duration=duration, seed=seed, steps=steps, strength=strength
        )
-        response_api = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/stability/v2beta/audio/stable-audio-2/audio-to-audio", method="POST"),
-            response_model=StabilityAudioResponse,
-            data=payload,
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/audio/stable-audio-2/audio-to-audio",
+                method=HttpMethod.POST,
+                request_model=StabilityAudioToAudioRequest,
+                response_model=StabilityAudioResponse,
+            ),
+            request=payload,
            content_type="multipart/form-data",
            files={"audio": audio_input_to_mp3(audio)},
+            auth_kwargs= {
+                "auth_token": cls.hidden.auth_token_comfy_org,
+                "comfy_api_key": cls.hidden.api_key_comfy_org,
+            },
        )
+        response_api = await operation.execute()
        if not response_api.audio:
            raise ValueError("No audio file was received in response.")
        return IO.NodeOutput(audio_bytes_to_audio_input(base64.b64decode(response_api.audio)))
@@ -856,14 +935,22 @@ class StabilityAudioInpaint(IO.ComfyNode):
            mask_start=mask_start,
            mask_end=mask_end,
        )
-        response_api = await sync_op(
-            cls,
-            endpoint=ApiEndpoint(path="/proxy/stability/v2beta/audio/stable-audio-2/inpaint", method="POST"),
-            response_model=StabilityAudioResponse,
-            data=payload,
+        operation = SynchronousOperation(
+            endpoint=ApiEndpoint(
+                path="/proxy/stability/v2beta/audio/stable-audio-2/inpaint",
+                method=HttpMethod.POST,
+                request_model=StabilityAudioInpaintRequest,
+                response_model=StabilityAudioResponse,
+            ),
+            request=payload,
            content_type="multipart/form-data",
            files={"audio": audio_input_to_mp3(audio)},
+            auth_kwargs={
+                "auth_token": cls.hidden.auth_token_comfy_org,
+                "comfy_api_key": cls.hidden.api_key_comfy_org,
+            },
        )
+        response_api = await operation.execute()
        if not response_api.audio:
            raise ValueError("No audio file was received in response.")
        return IO.NodeOutput(audio_bytes_to_audio_input(base64.b64decode(response_api.audio)))
--- a/comfy_api_nodes/nodes_topaz.py
+++ b/comfy_api_nodes/nodes_topaz.py
@@ -1,418 +0,0 @@
-import builtins
-from io import BytesIO
-
-import aiohttp
-import torch
-from typing_extensions import override
-
-from comfy_api.latest import IO, ComfyExtension, Input
-from comfy_api_nodes.apis import topaz_api
-from comfy_api_nodes.util import (
-    ApiEndpoint,
-    download_url_to_image_tensor,
-    download_url_to_video_output,
-    get_fs_object_size,
-    get_number_of_images,
-    poll_op,
-    sync_op,
-    upload_images_to_comfyapi,
-    validate_container_format_is_mp4,
-)
-
-UPSCALER_MODELS_MAP = {
-    "Starlight (Astra) Fast": "slf-1",
-    "Starlight (Astra) Creative": "slc-1",
-}
-UPSCALER_VALUES_MAP = {
-    "FullHD (1080p)": 1920,
-    "4K (2160p)": 3840,
-}
-
-
-class TopazImageEnhance(IO.ComfyNode):
-    @classmethod
-    def define_schema(cls):
-        return IO.Schema(
-            node_id="TopazImageEnhance",
-            display_name="Topaz Image Enhance",
-            category="api node/image/Topaz",
-            description="Industry-standard upscaling and image enhancement.",
-            inputs=[
-                IO.Combo.Input("model", options=["Reimagine"]),
-                IO.Image.Input("image"),
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    default="",
-                    tooltip="Optional text prompt for creative upscaling guidance.",
-                    optional=True,
-                ),
-                IO.Combo.Input(
-                    "subject_detection",
-                    options=["All", "Foreground", "Background"],
-                    optional=True,
-                ),
-                IO.Boolean.Input(
-                    "face_enhancement",
-                    default=True,
-                    optional=True,
-                    tooltip="Enhance faces (if present) during processing.",
-                ),
-                IO.Float.Input(
-                    "face_enhancement_creativity",
-                    default=0.0,
-                    min=0.0,
-                    max=1.0,
-                    step=0.01,
-                    display_mode=IO.NumberDisplay.number,
-                    optional=True,
-                    tooltip="Set the creativity level for face enhancement.",
-                ),
-                IO.Float.Input(
-                    "face_enhancement_strength",
-                    default=1.0,
-                    min=0.0,
-                    max=1.0,
-                    step=0.01,
-                    display_mode=IO.NumberDisplay.number,
-                    optional=True,
-                    tooltip="Controls how sharp enhanced faces are relative to the background.",
-                ),
-                IO.Boolean.Input(
-                    "crop_to_fill",
-                    default=False,
-                    optional=True,
-                    tooltip="By default, the image is letterboxed when the output aspect ratio differs. "
-                    "Enable to crop the image to fill the output dimensions.",
-                ),
-                IO.Int.Input(
-                    "output_width",
-                    default=0,
-                    min=0,
-                    max=32000,
-                    step=1,
-                    display_mode=IO.NumberDisplay.number,
-                    optional=True,
-                    tooltip="Zero value means to calculate automatically (usually it will be original size or output_height if specified).",
-                ),
-                IO.Int.Input(
-                    "output_height",
-                    default=0,
-                    min=0,
-                    max=32000,
-                    step=1,
-                    display_mode=IO.NumberDisplay.number,
-                    optional=True,
-                    tooltip="Zero value means to output in the same height as original or output width.",
-                ),
-                IO.Int.Input(
-                    "creativity",
-                    default=3,
-                    min=1,
-                    max=9,
-                    step=1,
-                    display_mode=IO.NumberDisplay.slider,
-                    optional=True,
-                ),
-                IO.Boolean.Input(
-                    "face_preservation",
-                    default=True,
-                    optional=True,
-                    tooltip="Preserve subjects' facial identity.",
-                ),
-                IO.Boolean.Input(
-                    "color_preservation",
-                    default=True,
-                    optional=True,
-                    tooltip="Preserve the original colors.",
-                ),
-            ],
-            outputs=[
-                IO.Image.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        model: str,
-        image: torch.Tensor,
-        prompt: str = "",
-        subject_detection: str = "All",
-        face_enhancement: bool = True,
-        face_enhancement_creativity: float = 1.0,
-        face_enhancement_strength: float = 0.8,
-        crop_to_fill: bool = False,
-        output_width: int = 0,
-        output_height: int = 0,
-        creativity: int = 3,
-        face_preservation: bool = True,
-        color_preservation: bool = True,
-    ) -> IO.NodeOutput:
-        if get_number_of_images(image) != 1:
-            raise ValueError("Only one input image is supported.")
-        download_url = await upload_images_to_comfyapi(cls, image, max_images=1, mime_type="image/png")
-        initial_response = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/topaz/image/v1/enhance-gen/async", method="POST"),
-            response_model=topaz_api.ImageAsyncTaskResponse,
-            data=topaz_api.ImageEnhanceRequest(
-                model=model,
-                prompt=prompt,
-                subject_detection=subject_detection,
-                face_enhancement=face_enhancement,
-                face_enhancement_creativity=face_enhancement_creativity,
-                face_enhancement_strength=face_enhancement_strength,
-                crop_to_fill=crop_to_fill,
-                output_width=output_width if output_width else None,
-                output_height=output_height if output_height else None,
-                creativity=creativity,
-                face_preservation=str(face_preservation).lower(),
-                color_preservation=str(color_preservation).lower(),
-                source_url=download_url[0],
-                output_format="png",
-            ),
-            content_type="multipart/form-data",
-        )
-
-        await poll_op(
-            cls,
-            poll_endpoint=ApiEndpoint(path=f"/proxy/topaz/image/v1/status/{initial_response.process_id}"),
-            response_model=topaz_api.ImageStatusResponse,
-            status_extractor=lambda x: x.status,
-            progress_extractor=lambda x: getattr(x, "progress", 0),
-            price_extractor=lambda x: x.credits * 0.08,
-            poll_interval=8.0,
-            max_poll_attempts=160,
-            estimated_duration=60,
-        )
-
-        results = await sync_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/topaz/image/v1/download/{initial_response.process_id}"),
-            response_model=topaz_api.ImageDownloadResponse,
-            monitor_progress=False,
-        )
-        return IO.NodeOutput(await download_url_to_image_tensor(results.download_url))
-
-
-class TopazVideoEnhance(IO.ComfyNode):
-    @classmethod
-    def define_schema(cls):
-        return IO.Schema(
-            node_id="TopazVideoEnhance",
-            display_name="Topaz Video Enhance",
-            category="api node/video/Topaz",
-            description="Breathe new life into video with powerful upscaling and recovery technology.",
-            inputs=[
-                IO.Video.Input("video"),
-                IO.Boolean.Input("upscaler_enabled", default=True),
-                IO.Combo.Input("upscaler_model", options=list(UPSCALER_MODELS_MAP.keys())),
-                IO.Combo.Input("upscaler_resolution", options=list(UPSCALER_VALUES_MAP.keys())),
-                IO.Combo.Input(
-                    "upscaler_creativity",
-                    options=["low", "middle", "high"],
-                    default="low",
-                    tooltip="Creativity level (applies only to Starlight (Astra) Creative).",
-                    optional=True,
-                ),
-                IO.Boolean.Input("interpolation_enabled", default=False, optional=True),
-                IO.Combo.Input("interpolation_model", options=["apo-8"], default="apo-8", optional=True),
-                IO.Int.Input(
-                    "interpolation_slowmo",
-                    default=1,
-                    min=1,
-                    max=16,
-                    display_mode=IO.NumberDisplay.number,
-                    tooltip="Slow-motion factor applied to the input video. "
-                    "For example, 2 makes the output twice as slow and doubles the duration.",
-                    optional=True,
-                ),
-                IO.Int.Input(
-                    "interpolation_frame_rate",
-                    default=60,
-                    min=15,
-                    max=240,
-                    display_mode=IO.NumberDisplay.number,
-                    tooltip="Output frame rate.",
-                    optional=True,
-                ),
-                IO.Boolean.Input(
-                    "interpolation_duplicate",
-                    default=False,
-                    tooltip="Analyze the input for duplicate frames and remove them.",
-                    optional=True,
-                ),
-                IO.Float.Input(
-                    "interpolation_duplicate_threshold",
-                    default=0.01,
-                    min=0.001,
-                    max=0.1,
-                    step=0.001,
-                    display_mode=IO.NumberDisplay.number,
-                    tooltip="Detection sensitivity for duplicate frames.",
-                    optional=True,
-                ),
-                IO.Combo.Input(
-                    "dynamic_compression_level",
-                    options=["Low", "Mid", "High"],
-                    default="Low",
-                    tooltip="CQP level.",
-                    optional=True,
-                ),
-            ],
-            outputs=[
-                IO.Video.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        video: Input.Video,
-        upscaler_enabled: bool,
-        upscaler_model: str,
-        upscaler_resolution: str,
-        upscaler_creativity: str = "low",
-        interpolation_enabled: bool = False,
-        interpolation_model: str = "apo-8",
-        interpolation_slowmo: int = 1,
-        interpolation_frame_rate: int = 60,
-        interpolation_duplicate: bool = False,
-        interpolation_duplicate_threshold: float = 0.01,
-        dynamic_compression_level: str = "Low",
-    ) -> IO.NodeOutput:
-        if upscaler_enabled is False and interpolation_enabled is False:
-            raise ValueError("There is nothing to do: both upscaling and interpolation are disabled.")
-        validate_container_format_is_mp4(video)
-        src_width, src_height = video.get_dimensions()
-        src_frame_rate = int(video.get_frame_rate())
-        duration_sec = video.get_duration()
-        src_video_stream = video.get_stream_source()
-        target_width = src_width
-        target_height = src_height
-        target_frame_rate = src_frame_rate
-        filters = []
-        if upscaler_enabled:
-            target_width = UPSCALER_VALUES_MAP[upscaler_resolution]
-            target_height = UPSCALER_VALUES_MAP[upscaler_resolution]
-            filters.append(
-                topaz_api.VideoEnhancementFilter(
-                    model=UPSCALER_MODELS_MAP[upscaler_model],
-                    creativity=(upscaler_creativity if UPSCALER_MODELS_MAP[upscaler_model] == "slc-1" else None),
-                    isOptimizedMode=(True if UPSCALER_MODELS_MAP[upscaler_model] == "slc-1" else None),
-                ),
-            )
-        if interpolation_enabled:
-            target_frame_rate = interpolation_frame_rate
-            filters.append(
-                topaz_api.VideoFrameInterpolationFilter(
-                    model=interpolation_model,
-                    slowmo=interpolation_slowmo,
-                    fps=interpolation_frame_rate,
-                    duplicate=interpolation_duplicate,
-                    duplicate_threshold=interpolation_duplicate_threshold,
-                ),
-            )
-        initial_res = await sync_op(
-            cls,
-            ApiEndpoint(path="/proxy/topaz/video/", method="POST"),
-            response_model=topaz_api.CreateVideoResponse,
-            data=topaz_api.CreateVideoRequest(
-                source=topaz_api.CreateCreateVideoRequestSource(
-                    container="mp4",
-                    size=get_fs_object_size(src_video_stream),
-                    duration=int(duration_sec),
-                    frameCount=video.get_frame_count(),
-                    frameRate=src_frame_rate,
-                    resolution=topaz_api.Resolution(width=src_width, height=src_height),
-                ),
-                filters=filters,
-                output=topaz_api.OutputInformationVideo(
-                    resolution=topaz_api.Resolution(width=target_width, height=target_height),
-                    frameRate=target_frame_rate,
-                    audioCodec="AAC",
-                    audioTransfer="Copy",
-                    dynamicCompressionLevel=dynamic_compression_level,
-                ),
-            ),
-            wait_label="Creating task",
-            final_label_on_success="Task created",
-        )
-        upload_res = await sync_op(
-            cls,
-            ApiEndpoint(
-                path=f"/proxy/topaz/video/{initial_res.requestId}/accept",
-                method="PATCH",
-            ),
-            response_model=topaz_api.VideoAcceptResponse,
-            wait_label="Preparing upload",
-            final_label_on_success="Upload started",
-        )
-        if len(upload_res.urls) > 1:
-            raise NotImplementedError(
-                "Large files are not currently supported. Please open an issue in the ComfyUI repository."
-            )
-        async with aiohttp.ClientSession(headers={"Content-Type": "video/mp4"}) as session:
-            if isinstance(src_video_stream, BytesIO):
-                src_video_stream.seek(0)
-                async with session.put(upload_res.urls[0], data=src_video_stream, raise_for_status=True) as res:
-                    upload_etag = res.headers["Etag"]
-            else:
-                with builtins.open(src_video_stream, "rb") as video_file:
-                    async with session.put(upload_res.urls[0], data=video_file, raise_for_status=True) as res:
-                        upload_etag = res.headers["Etag"]
-        await sync_op(
-            cls,
-            ApiEndpoint(
-                path=f"/proxy/topaz/video/{initial_res.requestId}/complete-upload",
-                method="PATCH",
-            ),
-            response_model=topaz_api.VideoCompleteUploadResponse,
-            data=topaz_api.VideoCompleteUploadRequest(
-                uploadResults=[
-                    topaz_api.VideoCompleteUploadRequestPart(
-                        partNum=1,
-                        eTag=upload_etag,
-                    ),
-                ],
-            ),
-            wait_label="Finalizing upload",
-            final_label_on_success="Upload completed",
-        )
-        final_response = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/topaz/video/{initial_res.requestId}/status"),
-            response_model=topaz_api.VideoStatusResponse,
-            status_extractor=lambda x: x.status,
-            progress_extractor=lambda x: getattr(x, "progress", 0),
-            price_extractor=lambda x: (x.estimates.cost[0] * 0.08 if x.estimates and x.estimates.cost[0] else None),
-            poll_interval=10.0,
-            max_poll_attempts=320,
-        )
-        return IO.NodeOutput(await download_url_to_video_output(final_response.download.url))
-
-
-class TopazExtension(ComfyExtension):
-    @override
-    async def get_node_list(self) -> list[type[IO.ComfyNode]]:
-        return [
-            TopazImageEnhance,
-            TopazVideoEnhance,
-        ]
-
-
-async def comfy_entrypoint() -> TopazExtension:
-    return TopazExtension()
--- a/comfy_api_nodes/nodes_veo2.py
+++ b/comfy_api_nodes/nodes_veo2.py
@@ -1,7 +1,6 @@
 import base64
 from io import BytesIO

-import torch
 from typing_extensions import override

 from comfy_api.input_impl.video_types import VideoFromFile
@@ -11,9 +10,6 @@ from comfy_api_nodes.apis.veo_api import (
    VeoGenVidPollResponse,
    VeoGenVidRequest,
    VeoGenVidResponse,
-    VeoRequestInstance,
-    VeoRequestInstanceImage,
-    VeoRequestParameters,
 )
 from comfy_api_nodes.util import (
    ApiEndpoint,
@@ -350,163 +346,12 @@ class Veo3VideoGenerationNode(VeoVideoGenerationNode):
        )


-class Veo3FirstLastFrameNode(IO.ComfyNode):
-
-    @classmethod
-    def define_schema(cls):
-        return IO.Schema(
-            node_id="Veo3FirstLastFrameNode",
-            display_name="Google Veo 3 First-Last-Frame to Video",
-            category="api node/video/Veo",
-            description="Generate video using prompt and first and last frames.",
-            inputs=[
-                IO.String.Input(
-                    "prompt",
-                    multiline=True,
-                    default="",
-                    tooltip="Text description of the video",
-                ),
-                IO.String.Input(
-                    "negative_prompt",
-                    multiline=True,
-                    default="",
-                    tooltip="Negative text prompt to guide what to avoid in the video",
-                ),
-                IO.Combo.Input("resolution", options=["720p", "1080p"]),
-                IO.Combo.Input(
-                    "aspect_ratio",
-                    options=["16:9", "9:16"],
-                    default="16:9",
-                    tooltip="Aspect ratio of the output video",
-                ),
-                IO.Int.Input(
-                    "duration",
-                    default=8,
-                    min=4,
-                    max=8,
-                    step=2,
-                    display_mode=IO.NumberDisplay.slider,
-                    tooltip="Duration of the output video in seconds",
-                ),
-                IO.Int.Input(
-                    "seed",
-                    default=0,
-                    min=0,
-                    max=0xFFFFFFFF,
-                    step=1,
-                    display_mode=IO.NumberDisplay.number,
-                    control_after_generate=True,
-                    tooltip="Seed for video generation",
-                ),
-                IO.Image.Input("first_frame", tooltip="Start frame"),
-                IO.Image.Input("last_frame", tooltip="End frame"),
-                IO.Combo.Input(
-                    "model",
-                    options=["veo-3.1-generate", "veo-3.1-fast-generate"],
-                    default="veo-3.1-fast-generate",
-                ),
-                IO.Boolean.Input(
-                    "generate_audio",
-                    default=True,
-                    tooltip="Generate audio for the video.",
-                ),
-            ],
-            outputs=[
-                IO.Video.Output(),
-            ],
-            hidden=[
-                IO.Hidden.auth_token_comfy_org,
-                IO.Hidden.api_key_comfy_org,
-                IO.Hidden.unique_id,
-            ],
-            is_api_node=True,
-        )
-
-    @classmethod
-    async def execute(
-        cls,
-        prompt: str,
-        negative_prompt: str,
-        resolution: str,
-        aspect_ratio: str,
-        duration: int,
-        seed: int,
-        first_frame: torch.Tensor,
-        last_frame: torch.Tensor,
-        model: str,
-        generate_audio: bool,
-    ):
-        model = MODELS_MAP[model]
-        initial_response = await sync_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/veo/{model}/generate", method="POST"),
-            response_model=VeoGenVidResponse,
-            data=VeoGenVidRequest(
-                instances=[
-                    VeoRequestInstance(
-                        prompt=prompt,
-                        image=VeoRequestInstanceImage(
-                            bytesBase64Encoded=tensor_to_base64_string(first_frame), mimeType="image/png"
-                        ),
-                        lastFrame=VeoRequestInstanceImage(
-                            bytesBase64Encoded=tensor_to_base64_string(last_frame), mimeType="image/png"
-                        ),
-                    ),
-                ],
-                parameters=VeoRequestParameters(
-                    aspectRatio=aspect_ratio,
-                    personGeneration="ALLOW",
-                    durationSeconds=duration,
-                    enhancePrompt=True,  # cannot be False for Veo3
-                    seed=seed,
-                    generateAudio=generate_audio,
-                    negativePrompt=negative_prompt,
-                    resolution=resolution,
-                ),
-            ),
-        )
-        poll_response = await poll_op(
-            cls,
-            ApiEndpoint(path=f"/proxy/veo/{model}/poll", method="POST"),
-            response_model=VeoGenVidPollResponse,
-            status_extractor=lambda r: "completed" if r.done else "pending",
-            data=VeoGenVidPollRequest(
-                operationName=initial_response.name,
-            ),
-            poll_interval=5.0,
-            estimated_duration=AVERAGE_DURATION_VIDEO_GEN,
-        )
-
-        if poll_response.error:
-            raise Exception(f"Veo API error: {poll_response.error.message} (code: {poll_response.error.code})")
-
-        response = poll_response.response
-        filtered_count = response.raiMediaFilteredCount
-        if filtered_count:
-            reasons = response.raiMediaFilteredReasons or []
-            reason_part = f": {reasons[0]}" if reasons else ""
-            raise Exception(
-                f"Content blocked by Google's Responsible AI filters{reason_part} "
-                f"({filtered_count} video{'s' if filtered_count != 1 else ''} filtered)."
-            )
-
-        if response.videos:
-            video = response.videos[0]
-            if video.bytesBase64Encoded:
-                return IO.NodeOutput(VideoFromFile(BytesIO(base64.b64decode(video.bytesBase64Encoded))))
-            if video.gcsUri:
-                return IO.NodeOutput(await download_url_to_video_output(video.gcsUri))
-            raise Exception("Video returned but no data or URL was provided")
-        raise Exception("Video generation completed but no video was returned")
-
-
 class VeoExtension(ComfyExtension):
    @override
    async def get_node_list(self) -> list[type[IO.ComfyNode]]:
        return [
            VeoVideoGenerationNode,
            Veo3VideoGenerationNode,
-            Veo3FirstLastFrameNode,
        ]


--- a/comfy_api_nodes/nodes_vidu.py
+++ b/comfy_api_nodes/nodes_vidu.py
@@ -14,9 +14,9 @@ from comfy_api_nodes.util import (
    poll_op,
    sync_op,
    upload_images_to_comfyapi,
-    validate_image_aspect_ratio,
+    validate_aspect_ratio_closeness,
+    validate_image_aspect_ratio_range,
    validate_image_dimensions,
-    validate_images_aspect_ratio_closeness,
 )

 VIDU_TEXT_TO_VIDEO = "/proxy/vidu/text2video"
@@ -114,7 +114,7 @@ async def execute_task(
        cls,
        ApiEndpoint(path=VIDU_GET_GENERATION_STATUS % response.task_id),
        response_model=TaskStatusResponse,
-        status_extractor=lambda r: r.state,
+        status_extractor=lambda r: r.state.value,
        estimated_duration=estimated_duration,
    )

@@ -307,7 +307,7 @@ class ViduImageToVideoNode(IO.ComfyNode):
    ) -> IO.NodeOutput:
        if get_number_of_images(image) > 1:
            raise ValueError("Only one input image is allowed.")
-        validate_image_aspect_ratio(image, (1, 4), (4, 1))
+        validate_image_aspect_ratio_range(image, (1, 4), (4, 1))
        payload = TaskCreationRequest(
            model_name=model,
            prompt=prompt,
@@ -423,7 +423,7 @@ class ViduReferenceVideoNode(IO.ComfyNode):
        if a > 7:
            raise ValueError("Too many images, maximum allowed is 7.")
        for image in images:
-            validate_image_aspect_ratio(image, (1, 4), (4, 1))
+            validate_image_aspect_ratio_range(image, (1, 4), (4, 1))
            validate_image_dimensions(image, min_width=128, min_height=128)
        payload = TaskCreationRequest(
            model_name=model,
@@ -533,7 +533,7 @@ class ViduStartEndToVideoNode(IO.ComfyNode):
        resolution: str,
        movement_amplitude: str,
    ) -> IO.NodeOutput:
-        validate_images_aspect_ratio_closeness(first_frame, end_frame, min_rel=0.8, max_rel=1.25, strict=False)
+        validate_aspect_ratio_closeness(first_frame, end_frame, min_rel=0.8, max_rel=1.25, strict=False)
        payload = TaskCreationRequest(
            model_name=model,
            prompt=prompt,
--- a/comfy_api_nodes/util/init.py
+++ b/comfy_api_nodes/util/init.py
@@ -14,12 +14,9 @@ from .conversions import (
    downscale_image_tensor,
    image_tensor_pair_to_batch,
    pil_to_bytesio,
-    resize_mask_to_image,
    tensor_to_base64_string,
    tensor_to_bytesio,
    tensor_to_pil,
-    text_filepath_to_base64_string,
-    text_filepath_to_data_uri,
    trim_video,
    video_to_base64_string,
 )
@@ -36,14 +33,13 @@ from .upload_helpers import (
    upload_video_to_comfyapi,
 )
 from .validation_utils import (
-    get_image_dimensions,
    get_number_of_images,
-    validate_aspect_ratio_string,
+    validate_aspect_ratio_closeness,
    validate_audio_duration,
    validate_container_format_is_mp4,
    validate_image_aspect_ratio,
+    validate_image_aspect_ratio_range,
    validate_image_dimensions,
-    validate_images_aspect_ratio_closeness,
    validate_string,
    validate_video_dimensions,
    validate_video_duration,
@@ -74,23 +70,19 @@ __all__ = [
    "downscale_image_tensor",
    "image_tensor_pair_to_batch",
    "pil_to_bytesio",
-    "resize_mask_to_image",
    "tensor_to_base64_string",
    "tensor_to_bytesio",
    "tensor_to_pil",
-    "text_filepath_to_base64_string",
-    "text_filepath_to_data_uri",
    "trim_video",
    "video_to_base64_string",
    # Validation utilities
-    "get_image_dimensions",
    "get_number_of_images",
-    "validate_aspect_ratio_string",
+    "validate_aspect_ratio_closeness",
    "validate_audio_duration",
    "validate_container_format_is_mp4",
    "validate_image_aspect_ratio",
+    "validate_image_aspect_ratio_range",
    "validate_image_dimensions",
-    "validate_images_aspect_ratio_closeness",
    "validate_string",
    "validate_video_dimensions",
    "validate_video_duration",
--- a/comfy_api_nodes/util/client.py
+++ b/comfy_api_nodes/util/client.py
@@ -16,9 +16,9 @@ from pydantic import BaseModel

 from comfy import utils
 from comfy_api.latest import IO
+from comfy_api_nodes.apis import request_logger
 from server import PromptServer

-from . import request_logger
 from ._helpers import (
    default_base_url,
    get_auth_header,
@@ -63,7 +63,6 @@ class _RequestConfig:
    estimated_total: Optional[int] = None
    final_label_on_success: Optional[str] = "Completed"
    progress_origin_ts: Optional[float] = None
-    price_extractor: Optional[Callable[[dict[str, Any]], Optional[float]]] = None


@dataclass
@@ -78,9 +77,9 @@ class _PollUIState:


 _RETRY_STATUS = {408, 429, 500, 502, 503, 504}
-COMPLETED_STATUSES = ["succeeded", "succeed", "success", "completed", "finished", "done", "complete"]
-FAILED_STATUSES = ["cancelled", "canceled", "canceling", "fail", "failed", "error"]
-QUEUED_STATUSES = ["created", "queued", "queueing", "submitted", "initializing"]
+COMPLETED_STATUSES = ["succeeded", "succeed", "success", "completed"]
+FAILED_STATUSES = ["cancelled", "canceled", "failed", "error"]
+QUEUED_STATUSES = ["created", "queued", "queueing", "submitted"]


 async def sync_op(
@@ -88,7 +87,6 @@ async def sync_op(
    endpoint: ApiEndpoint,
    *,
    response_model: Type[M],
-    price_extractor: Optional[Callable[[M], Optional[float]]] = None,
    data: Optional[BaseModel] = None,
    files: Optional[Union[dict[str, Any], list[tuple[str, Any]]]] = None,
    content_type: str = "application/json",
@@ -106,7 +104,6 @@ async def sync_op(
    raw = await sync_op_raw(
        cls,
        endpoint,
-        price_extractor=_wrap_model_extractor(response_model, price_extractor),
        data=data,
        files=files,
        content_type=content_type,
@@ -178,7 +175,6 @@ async def sync_op_raw(
    cls: type[IO.ComfyNode],
    endpoint: ApiEndpoint,
    *,
-    price_extractor: Optional[Callable[[dict[str, Any]], Optional[float]]] = None,
    data: Optional[Union[dict[str, Any], BaseModel]] = None,
    files: Optional[Union[dict[str, Any], list[tuple[str, Any]]]] = None,
    content_type: str = "application/json",
@@ -220,7 +216,6 @@ async def sync_op_raw(
        estimated_total=estimated_duration,
        final_label_on_success=final_label_on_success,
        progress_origin_ts=progress_origin_ts,
-        price_extractor=price_extractor,
    )
    return await _request_base(cfg, expect_binary=as_binary)

@@ -429,9 +424,7 @@ def _display_text(
    if status:
        display_lines.append(f"Status: {status.capitalize() if isinstance(status, str) else status}")
    if price is not None:
-        p = f"{float(price):,.4f}".rstrip("0").rstrip(".")
-        if p != "0":
-            display_lines.append(f"Price: ${p}")
+        display_lines.append(f"Price: ${float(price):,.4f}")
    if text is not None:
        display_lines.append(text)
    if display_lines:
@@ -587,7 +580,6 @@ async def _request_base(cfg: _RequestConfig, expect_binary: bool):
    delay = cfg.retry_delay
    operation_succeeded: bool = False
    final_elapsed_seconds: Optional[int] = None
-    extracted_price: Optional[float] = None
    while True:
        attempt += 1
        stop_event = asyncio.Event()
@@ -597,7 +589,7 @@ async def _request_base(cfg: _RequestConfig, expect_binary: bool):
        operation_id = _generate_operation_id(method, cfg.endpoint.path, attempt)
        logging.debug("[DEBUG] HTTP %s %s (attempt %d)", method, url, attempt)

-        payload_headers = {"Accept": "*/*"} if expect_binary else {"Accept": "application/json"}
+        payload_headers = {"Accept": "*/*"}
        if not parsed_url.scheme and not parsed_url.netloc:  # is URL relative?
            payload_headers.update(get_auth_header(cfg.node_cls))
        if cfg.endpoint.headers:
@@ -775,8 +767,6 @@ async def _request_base(cfg: _RequestConfig, expect_binary: bool):
                        except json.JSONDecodeError:
                            payload = {"_raw": text}
                        response_content_to_log = payload if isinstance(payload, dict) else text
-                    with contextlib.suppress(Exception):
-                        extracted_price = cfg.price_extractor(payload) if cfg.price_extractor else None
                    operation_succeeded = True
                    final_elapsed_seconds = int(time.monotonic() - start_time)
                    try:
@@ -881,7 +871,7 @@ async def _request_base(cfg: _RequestConfig, expect_binary: bool):
                        else int(time.monotonic() - start_time)
                    ),
                    estimated_total=cfg.estimated_total,
-                    price=extracted_price,
+                    price=None,
                    is_queued=False,
                    processing_elapsed_seconds=final_elapsed_seconds,
                )
--- a/comfy_api_nodes/util/conversions.py
+++ b/comfy_api_nodes/util/conversions.py
@@ -1,7 +1,6 @@
 import base64
 import logging
 import math
-import mimetypes
 import uuid
 from io import BytesIO
 from typing import Optional
@@ -13,7 +12,7 @@ from PIL import Image

 from comfy.utils import common_upscale
 from comfy_api.latest import Input, InputImpl
-from comfy_api.util import VideoCodec, VideoContainer
+from comfy_api.util import VideoContainer, VideoCodec

 from ._helpers import mimetype_to_extension

@@ -431,40 +430,3 @@ def audio_bytes_to_audio_input(audio_bytes: bytes) -> dict:
    wav = torch.cat(frames, dim=1)  # [C, T]
    wav = _f32_pcm(wav)
    return {"waveform": wav.unsqueeze(0).contiguous(), "sample_rate": out_sr}
-
-
-def resize_mask_to_image(
-    mask: torch.Tensor,
-    image: torch.Tensor,
-    upscale_method="nearest-exact",
-    crop="disabled",
-    allow_gradient=True,
-    add_channel_dim=False,
-):
-    """Resize mask to be the same dimensions as an image, while maintaining proper format for API calls."""
-    _, height, width, _ = image.shape
-    mask = mask.unsqueeze(-1)
-    mask = mask.movedim(-1, 1)
-    mask = common_upscale(mask, width=width, height=height, upscale_method=upscale_method, crop=crop)
-    mask = mask.movedim(1, -1)
-    if not add_channel_dim:
-        mask = mask.squeeze(-1)
-    if not allow_gradient:
-        mask = (mask > 0.5).float()
-    return mask
-
-
-def text_filepath_to_base64_string(filepath: str) -> str:
-    """Converts a text file to a base64 string."""
-    with open(filepath, "rb") as f:
-        file_content = f.read()
-    return base64.b64encode(file_content).decode("utf-8")
-
-
-def text_filepath_to_data_uri(filepath: str) -> str:
-    """Converts a text file to a data URI."""
-    base64_string = text_filepath_to_base64_string(filepath)
-    mime_type, _ = mimetypes.guess_type(filepath)
-    if mime_type is None:
-        mime_type = "application/octet-stream"
-    return f"data:{mime_type};base64,{base64_string}"
--- a/comfy_api_nodes/util/download_helpers.py
+++ b/comfy_api_nodes/util/download_helpers.py
@@ -12,8 +12,8 @@ from aiohttp.client_exceptions import ClientError, ContentTypeError

 from comfy_api.input_impl import VideoFromFile
 from comfy_api.latest import IO as COMFY_IO
+from comfy_api_nodes.apis import request_logger

-from . import request_logger
 from ._helpers import (
    default_base_url,
    get_auth_header,
@@ -232,12 +232,11 @@ async def download_url_to_video_output(
    video_url: str,
    *,
    timeout: float = None,
-    max_retries: int = 5,
    cls: type[COMFY_IO.ComfyNode] = None,
 ) -> VideoFromFile:
    """Downloads a video from a URL and returns a `VIDEO` output."""
    result = BytesIO()
-    await download_url_to_bytesio(video_url, result, timeout=timeout, max_retries=max_retries, cls=cls)
+    await download_url_to_bytesio(video_url, result, timeout=timeout, cls=cls)
    return VideoFromFile(result)


--- a/comfy_api_nodes/util/upload_helpers.py
+++ b/comfy_api_nodes/util/upload_helpers.py
@@ -4,7 +4,7 @@ import logging
 import time
 import uuid
 from io import BytesIO
-from typing import Optional
+from typing import Optional, Union
 from urllib.parse import urlparse

 import aiohttp
@@ -13,8 +13,8 @@ from pydantic import BaseModel, Field

 from comfy_api.latest import IO, Input
 from comfy_api.util import VideoCodec, VideoContainer
+from comfy_api_nodes.apis import request_logger

-from . import request_logger
 from ._helpers import is_processing_interrupted, sleep_with_interrupt
 from .client import (
    ApiEndpoint,
@@ -48,9 +48,8 @@ async def upload_images_to_comfyapi(
    image: torch.Tensor,
    *,
    max_images: int = 8,
-    mime_type: str | None = None,
-    wait_label: str | None = "Uploading",
-    show_batch_index: bool = True,
+    mime_type: Optional[str] = None,
+    wait_label: Optional[str] = "Uploading",
 ) -> list[str]:
    """
    Uploads images to ComfyUI API and returns download URLs.
@@ -60,18 +59,11 @@ async def upload_images_to_comfyapi(
    download_urls: list[str] = []
    is_batch = len(image.shape) > 3
    batch_len = image.shape[0] if is_batch else 1
-    num_to_upload = min(batch_len, max_images)
-    batch_start_ts = time.monotonic()

-    for idx in range(num_to_upload):
+    for idx in range(min(batch_len, max_images)):
        tensor = image[idx] if is_batch else image
        img_io = tensor_to_bytesio(tensor, mime_type=mime_type)
-
-        effective_label = wait_label
-        if wait_label and show_batch_index and num_to_upload > 1:
-            effective_label = f"{wait_label} ({idx + 1}/{num_to_upload})"
-
-        url = await upload_file_to_comfyapi(cls, img_io, img_io.name, mime_type, effective_label, batch_start_ts)
+        url = await upload_file_to_comfyapi(cls, img_io, img_io.name, mime_type, wait_label)
        download_urls.append(url)
    return download_urls

@@ -103,7 +95,6 @@ async def upload_video_to_comfyapi(
    container: VideoContainer = VideoContainer.MP4,
    codec: VideoCodec = VideoCodec.H264,
    max_duration: Optional[int] = None,
-    wait_label: str | None = "Uploading",
 ) -> str:
    """
    Uploads a single video to ComfyUI API and returns its download URL.
@@ -128,16 +119,15 @@ async def upload_video_to_comfyapi(
    video.save_to(video_bytes_io, format=container, codec=codec)
    video_bytes_io.seek(0)

-    return await upload_file_to_comfyapi(cls, video_bytes_io, filename, upload_mime_type, wait_label)
+    return await upload_file_to_comfyapi(cls, video_bytes_io, filename, upload_mime_type)


 async def upload_file_to_comfyapi(
    cls: type[IO.ComfyNode],
    file_bytes_io: BytesIO,
    filename: str,
-    upload_mime_type: str | None,
-    wait_label: str | None = "Uploading",
-    progress_origin_ts: float | None = None,
+    upload_mime_type: Optional[str],
+    wait_label: Optional[str] = "Uploading",
 ) -> str:
    """Uploads a single file to ComfyUI API and returns its download URL."""
    if upload_mime_type is None:
@@ -158,7 +148,6 @@ async def upload_file_to_comfyapi(
        file_bytes_io,
        content_type=upload_mime_type,
        wait_label=wait_label,
-        progress_origin_ts=progress_origin_ts,
    )
    return create_resp.download_url

@@ -166,18 +155,27 @@ async def upload_file_to_comfyapi(
 async def upload_file(
    cls: type[IO.ComfyNode],
    upload_url: str,
-    file: BytesIO | str,
+    file: Union[BytesIO, str],
    *,
-    content_type: str | None = None,
+    content_type: Optional[str] = None,
    max_retries: int = 3,
    retry_delay: float = 1.0,
    retry_backoff: float = 2.0,
-    wait_label: str | None = None,
-    progress_origin_ts: float | None = None,
+    wait_label: Optional[str] = None,
 ) -> None:
    """
    Upload a file to a signed URL (e.g., S3 pre-signed PUT) with retries, Comfy progress display, and interruption.

+    Args:
+        cls: Node class (provides auth context + UI progress hooks).
+        upload_url: Pre-signed PUT URL.
+        file: BytesIO or path string.
+        content_type: Explicit MIME type. If None, we *suppress* Content-Type.
+        max_retries: Maximum retry attempts.
+        retry_delay: Initial delay in seconds.
+        retry_backoff: Exponential backoff factor.
+        wait_label: Progress label shown in Comfy UI.
+
    Raises:
        ProcessingInterrupted, LocalNetworkError, ApiServerError, Exception
    """
@@ -200,7 +198,7 @@ async def upload_file(

    attempt = 0
    delay = retry_delay
-    start_ts = progress_origin_ts if progress_origin_ts is not None else time.monotonic()
+    start_ts = time.monotonic()
    op_uuid = uuid.uuid4().hex[:8]
    while True:
        attempt += 1
--- a/comfy_api_nodes/util/validation_utils.py
+++ b/comfy_api_nodes/util/validation_utils.py
@@ -37,62 +37,63 @@ def validate_image_dimensions(

 def validate_image_aspect_ratio(
    image: torch.Tensor,
-    min_ratio: Optional[tuple[float, float]] = None,  # e.g. (1, 4)
-    max_ratio: Optional[tuple[float, float]] = None,  # e.g. (4, 1)
+    min_aspect_ratio: Optional[float] = None,
+    max_aspect_ratio: Optional[float] = None,
+):
+    width, height = get_image_dimensions(image)
+    aspect_ratio = width / height
+
+    if min_aspect_ratio is not None and aspect_ratio < min_aspect_ratio:
+        raise ValueError(f"Image aspect ratio must be at least {min_aspect_ratio}, got {aspect_ratio}")
+    if max_aspect_ratio is not None and aspect_ratio > max_aspect_ratio:
+        raise ValueError(f"Image aspect ratio must be at most {max_aspect_ratio}, got {aspect_ratio}")
+
+
+def validate_image_aspect_ratio_range(
+    image: torch.Tensor,
+    min_ratio: tuple[float, float],  # e.g. (1, 4)
+    max_ratio: tuple[float, float],  # e.g. (4, 1)
    *,
    strict: bool = True,  # True -> (min, max); False -> [min, max]
 ) -> float:
-    """Validates that image aspect ratio is within min and max. If a bound is None, that side is not checked."""
+    a1, b1 = min_ratio
+    a2, b2 = max_ratio
+    if a1 <= 0 or b1 <= 0 or a2 <= 0 or b2 <= 0:
+        raise ValueError("Ratios must be positive, like (1, 4) or (4, 1).")
+    lo, hi = (a1 / b1), (a2 / b2)
+    if lo > hi:
+        lo, hi = hi, lo
+        a1, b1, a2, b2 = a2, b2, a1, b1  # swap only for error text
    w, h = get_image_dimensions(image)
    if w <= 0 or h <= 0:
        raise ValueError(f"Invalid image dimensions: {w}x{h}")
    ar = w / h
-    _assert_ratio_bounds(ar, min_ratio=min_ratio, max_ratio=max_ratio, strict=strict)
+    ok = (lo < ar < hi) if strict else (lo <= ar <= hi)
+    if not ok:
+        op = "<" if strict else "≤"
+        raise ValueError(f"Image aspect ratio {ar:.6g} is outside allowed range: {a1}:{b1} {op} ratio {op} {a2}:{b2}")
    return ar


-def validate_images_aspect_ratio_closeness(
-    first_image: torch.Tensor,
-    second_image: torch.Tensor,
-    min_rel: float,   # e.g. 0.8
-    max_rel: float,   # e.g. 1.25
+def validate_aspect_ratio_closeness(
+    start_img,
+    end_img,
+    min_rel: float,
+    max_rel: float,
    *,
-    strict: bool = False,  # True -> (min, max); False -> [min, max]
-) -> float:
-    """
-    Validates that the two images' aspect ratios are 'close'.
-    The closeness factor is C = max(ar1, ar2) / min(ar1, ar2)  (C >= 1).
-    We require C <= limit, where limit = max(max_rel, 1.0 / min_rel).
-
-    Returns the computed closeness factor C.
-    """
-    w1, h1 = get_image_dimensions(first_image)
-    w2, h2 = get_image_dimensions(second_image)
+    strict: bool = False,  # True => exclusive, False => inclusive
+) -> None:
+    w1, h1 = get_image_dimensions(start_img)
+    w2, h2 = get_image_dimensions(end_img)
    if min(w1, h1, w2, h2) <= 0:
        raise ValueError("Invalid image dimensions")
    ar1 = w1 / h1
    ar2 = w2 / h2
+    # Normalize so it is symmetric (no need to check both ar1/ar2 and ar2/ar1)
    closeness = max(ar1, ar2) / min(ar1, ar2)
-    limit = max(max_rel, 1.0 / min_rel)
+    limit = max(max_rel, 1.0 / min_rel)  # for 0.8..1.25 this is 1.25
    if (closeness >= limit) if strict else (closeness > limit):
-        raise ValueError(
-            f"Aspect ratios must be close: ar1/ar2={ar1/ar2:.2g}, "
-            f"allowed range {min_rel}–{max_rel} (limit {limit:.2g})."
-        )
-    return closeness
-
-
-def validate_aspect_ratio_string(
-    aspect_ratio: str,
-    min_ratio: Optional[tuple[float, float]] = None,  # e.g. (1, 4)
-    max_ratio: Optional[tuple[float, float]] = None,  # e.g. (4, 1)
-    *,
-    strict: bool = False,  # True -> (min, max); False -> [min, max]
-) -> float:
-    """Parses 'X:Y' and validates it against optional bounds. Returns the numeric ratio."""
-    ar = _parse_aspect_ratio_string(aspect_ratio)
-    _assert_ratio_bounds(ar, min_ratio=min_ratio, max_ratio=max_ratio, strict=strict)
-    return ar
+        raise ValueError(f"Aspect ratios must be close: start/end={ar1/ar2:.4f}, allowed range {min_rel}–{max_rel}.")


 def validate_video_dimensions(
@@ -182,49 +183,3 @@ def validate_container_format_is_mp4(video: VideoInput) -> None:
    container_format = video.get_container_format()
    if container_format not in ["mp4", "mov,mp4,m4a,3gp,3g2,mj2"]:
        raise ValueError(f"Only MP4 container format supported. Got: {container_format}")
-
-
-def _ratio_from_tuple(r: tuple[float, float]) -> float:
-    a, b = r
-    if a <= 0 or b <= 0:
-        raise ValueError(f"Ratios must be positive, got {a}:{b}.")
-    return a / b
-
-
-def _assert_ratio_bounds(
-    ar: float,
-    *,
-    min_ratio: Optional[tuple[float, float]] = None,
-    max_ratio: Optional[tuple[float, float]] = None,
-    strict: bool = True,
-) -> None:
-    """Validate a numeric aspect ratio against optional min/max ratio bounds."""
-    lo = _ratio_from_tuple(min_ratio) if min_ratio is not None else None
-    hi = _ratio_from_tuple(max_ratio) if max_ratio is not None else None
-
-    if lo is not None and hi is not None and lo > hi:
-        lo, hi = hi, lo  # normalize order if caller swapped them
-
-    if lo is not None:
-        if (ar <= lo) if strict else (ar < lo):
-            op = "<" if strict else "≤"
-            raise ValueError(f"Aspect ratio `{ar:.2g}` must be {op} {lo:.2g}.")
-    if hi is not None:
-        if (ar >= hi) if strict else (ar > hi):
-            op = "<" if strict else "≤"
-            raise ValueError(f"Aspect ratio `{ar:.2g}` must be {op} {hi:.2g}.")
-
-
-def _parse_aspect_ratio_string(ar_str: str) -> float:
-    """Parse 'X:Y' with integer parts into a positive float ratio X/Y."""
-    parts = ar_str.split(":")
-    if len(parts) != 2:
-        raise ValueError(f"Aspect ratio must be 'X:Y' (e.g., 16:9), got '{ar_str}'.")
-    try:
-        a = int(parts[0].strip())
-        b = int(parts[1].strip())
-    except ValueError as exc:
-        raise ValueError(f"Aspect ratio must contain integers separated by ':', got '{ar_str}'.") from exc
-    if a <= 0 or b <= 0:
-        raise ValueError(f"Aspect ratio parts must be positive integers, got {a}:{b}.")
-    return a / b
--- a/comfy_execution/caching.py
+++ b/comfy_execution/caching.py
@@ -1,9 +1,4 @@
-import bisect
-import gc
 import itertools
-import psutil
-import time
-import torch
 from typing import Sequence, Mapping, Dict
 from comfy_execution.graph import DynamicPrompt
 from abc import ABC, abstractmethod
@@ -53,7 +48,7 @@ class Unhashable:
 def to_hashable(obj):
    # So that we don't infinitely recurse since frozenset and tuples
    # are Sequences.
-    if isinstance(obj, (int, float, str, bool, bytes, type(None))):
+    if isinstance(obj, (int, float, str, bool, type(None))):
        return obj
    elif isinstance(obj, Mapping):
        return frozenset([(to_hashable(k), to_hashable(v)) for k, v in sorted(obj.items())])
@@ -193,9 +188,6 @@ class BasicCache:
        self._clean_cache()
        self._clean_subcaches()

-    def poll(self, **kwargs):
-        pass
-
    def _set_immediate(self, node_id, value):
        assert self.initialized
        cache_key = self.cache_key_set.get_data_key(node_id)
@@ -284,9 +276,6 @@ class NullCache:
    def clean_unused(self):
        pass

-    def poll(self, **kwargs):
-        pass
-
    def get(self, node_id):
        return None

@@ -347,77 +336,3 @@ class LRUCache(BasicCache):
            self._mark_used(child_id)
            self.children[cache_key].append(self.cache_key_set.get_data_key(child_id))
        return self
-
-
-#Iterating the cache for usage analysis might be expensive, so if we trigger make sure
-#to take a chunk out to give breathing space on high-node / low-ram-per-node flows.
-
-RAM_CACHE_HYSTERESIS = 1.1
-
-#This is kinda in GB but not really. It needs to be non-zero for the below heuristic
-#and as long as Multi GB models dwarf this it will approximate OOM scoring OK
-
-RAM_CACHE_DEFAULT_RAM_USAGE = 0.1
-
-#Exponential bias towards evicting older workflows so garbage will be taken out
-#in constantly changing setups.
-
-RAM_CACHE_OLD_WORKFLOW_OOM_MULTIPLIER = 1.3
-
-class RAMPressureCache(LRUCache):
-
-    def __init__(self, key_class):
-        super().__init__(key_class, 0)
-        self.timestamps = {}
-
-    def clean_unused(self):
-        self._clean_subcaches()
-
-    def set(self, node_id, value):
-        self.timestamps[self.cache_key_set.get_data_key(node_id)] = time.time()
-        super().set(node_id, value)
-
-    def get(self, node_id):
-        self.timestamps[self.cache_key_set.get_data_key(node_id)] = time.time()
-        return super().get(node_id)
-
-    def poll(self, ram_headroom):
-        def _ram_gb():
-            return psutil.virtual_memory().available / (1024**3)
-
-        if _ram_gb() > ram_headroom:
-            return
-        gc.collect()
-        if _ram_gb() > ram_headroom:
-            return
-
-        clean_list = []
-
-        for key, (outputs, _), in self.cache.items():
-            oom_score =  RAM_CACHE_OLD_WORKFLOW_OOM_MULTIPLIER ** (self.generation - self.used_generation[key])
-
-            ram_usage = RAM_CACHE_DEFAULT_RAM_USAGE
-            def scan_list_for_ram_usage(outputs):
-                nonlocal ram_usage
-                if outputs is None:
-                    return
-                for output in outputs:
-                    if isinstance(output, list):
-                        scan_list_for_ram_usage(output)
-                    elif isinstance(output, torch.Tensor) and output.device.type == 'cpu':
-                        #score Tensors at a 50% discount for RAM usage as they are likely to
-                        #be high value intermediates
-                        ram_usage += (output.numel() * output.element_size()) * 0.5
-                    elif hasattr(output, "get_ram_usage"):
-                        ram_usage += output.get_ram_usage()
-            scan_list_for_ram_usage(outputs)
-
-            oom_score *= ram_usage
-            #In the case where we have no information on the node ram usage at all,
-            #break OOM score ties on the last touch timestamp (pure LRU)
-            bisect.insort(clean_list, (oom_score, self.timestamps[key], key))
-
-        while _ram_gb() < ram_headroom * RAM_CACHE_HYSTERESIS and clean_list:
-            _, _, key = clean_list.pop()
-            del self.cache[key]
-            gc.collect()
--- a/comfy_execution/graph.py
+++ b/comfy_execution/graph.py
@@ -209,15 +209,10 @@ class ExecutionList(TopologicalSort):
            self.execution_cache_listeners[from_node_id] = set()
        self.execution_cache_listeners[from_node_id].add(to_node_id)

-    def get_cache(self, from_node_id, to_node_id):
+    def get_output_cache(self, from_node_id, to_node_id):
        if not to_node_id in self.execution_cache:
            return None
-        value = self.execution_cache[to_node_id].get(from_node_id)
-        if value is None:
-            return None
-        #Write back to the main cache on touch.
-        self.output_cache.set(from_node_id, value)
-        return value
+        return self.execution_cache[to_node_id].get(from_node_id)

    def cache_update(self, node_id, value):
        if node_id in self.execution_cache_listeners:
--- a/comfy_extras/nodes_custom_sampler.py
+++ b/comfy_extras/nodes_custom_sampler.py
--- a/comfy_extras/nodes_dataset.py
+++ b/comfy_extras/nodes_dataset.py
--- a/comfy_extras/nodes_easycache.py
+++ b/comfy_extras/nodes_easycache.py
@@ -11,13 +11,13 @@ if TYPE_CHECKING:

 def easycache_forward_wrapper(executor, *args, **kwargs):
    # get values from args
+    x: torch.Tensor = args[0]
    transformer_options: dict[str] = args[-1]
    if not isinstance(transformer_options, dict):
        transformer_options = kwargs.get("transformer_options")
        if not transformer_options:
            transformer_options = args[-2]
    easycache: EasyCacheHolder = transformer_options["easycache"]
-    x: torch.Tensor = args[0][:, :easycache.output_channels]
    sigmas = transformer_options["sigmas"]
    uuids = transformer_options["uuids"]
    if sigmas is not None and easycache.is_past_end_timestep(sigmas):
@@ -82,13 +82,13 @@ def easycache_forward_wrapper(executor, *args, **kwargs):

 def lazycache_predict_noise_wrapper(executor, *args, **kwargs):
    # get values from args
+    x: torch.Tensor = args[0]
    timestep: float = args[1]
    model_options: dict[str] = args[2]
    easycache: LazyCacheHolder = model_options["transformer_options"]["easycache"]
    if easycache.is_past_end_timestep(timestep):
        return executor(*args, **kwargs)
    # prepare next x_prev
-    x: torch.Tensor = args[0][:, :easycache.output_channels]
    next_x_prev = x
    input_change = None
    do_easycache = easycache.should_do_easycache(timestep)
@@ -173,7 +173,7 @@ def easycache_sample_wrapper(executor, *args, **kwargs):


 class EasyCacheHolder:
-    def __init__(self, reuse_threshold: float, start_percent: float, end_percent: float, subsample_factor: int, offload_cache_diff: bool, verbose: bool=False, output_channels: int=None):
+    def __init__(self, reuse_threshold: float, start_percent: float, end_percent: float, subsample_factor: int, offload_cache_diff: bool, verbose: bool=False):
        self.name = "EasyCache"
        self.reuse_threshold = reuse_threshold
        self.start_percent = start_percent
@@ -202,7 +202,6 @@ class EasyCacheHolder:
        self.allow_mismatch = True
        self.cut_from_start = True
        self.state_metadata = None
-        self.output_channels = output_channels

    def is_past_end_timestep(self, timestep: float) -> bool:
        return not (timestep[0] > self.end_t).item()
@@ -265,7 +264,7 @@ class EasyCacheHolder:
                    else:
                        slicing.append(slice(None))
                batch_slice = batch_slice + slicing
-            x[tuple(batch_slice)] += self.uuid_cache_diffs[uuid].to(x.device)
+            x[batch_slice] += self.uuid_cache_diffs[uuid].to(x.device)
        return x

    def update_cache_diff(self, output: torch.Tensor, x: torch.Tensor, uuids: list[UUID]):
@@ -284,7 +283,7 @@ class EasyCacheHolder:
                else:
                    slicing.append(slice(None))
                skip_dim = False
-            x = x[tuple(slicing)]
+            x = x[slicing]
        diff = output - x
        batch_offset = diff.shape[0] // len(uuids)
        for i, uuid in enumerate(uuids):
@@ -324,7 +323,7 @@ class EasyCacheHolder:
        return self

    def clone(self):
-        return EasyCacheHolder(self.reuse_threshold, self.start_percent, self.end_percent, self.subsample_factor, self.offload_cache_diff, self.verbose, output_channels=self.output_channels)
+        return EasyCacheHolder(self.reuse_threshold, self.start_percent, self.end_percent, self.subsample_factor, self.offload_cache_diff, self.verbose)


 class EasyCacheNode(io.ComfyNode):
@@ -351,7 +350,7 @@ class EasyCacheNode(io.ComfyNode):
    @classmethod
    def execute(cls, model: io.Model.Type, reuse_threshold: float, start_percent: float, end_percent: float, verbose: bool) -> io.NodeOutput:
        model = model.clone()
-        model.model_options["transformer_options"]["easycache"] = EasyCacheHolder(reuse_threshold, start_percent, end_percent, subsample_factor=8, offload_cache_diff=False, verbose=verbose, output_channels=model.model.latent_format.latent_channels)
+        model.model_options["transformer_options"]["easycache"] = EasyCacheHolder(reuse_threshold, start_percent, end_percent, subsample_factor=8, offload_cache_diff=False, verbose=verbose)
        model.add_wrapper_with_key(comfy.patcher_extension.WrappersMP.OUTER_SAMPLE, "easycache", easycache_sample_wrapper)
        model.add_wrapper_with_key(comfy.patcher_extension.WrappersMP.CALC_COND_BATCH, "easycache", easycache_calc_cond_batch_wrapper)
        model.add_wrapper_with_key(comfy.patcher_extension.WrappersMP.DIFFUSION_MODEL, "easycache", easycache_forward_wrapper)
@@ -359,7 +358,7 @@ class EasyCacheNode(io.ComfyNode):


 class LazyCacheHolder:
-    def __init__(self, reuse_threshold: float, start_percent: float, end_percent: float, subsample_factor: int, offload_cache_diff: bool, verbose: bool=False, output_channels: int=None):
+    def __init__(self, reuse_threshold: float, start_percent: float, end_percent: float, subsample_factor: int, offload_cache_diff: bool, verbose: bool=False):
        self.name = "LazyCache"
        self.reuse_threshold = reuse_threshold
        self.start_percent = start_percent
@@ -383,7 +382,6 @@ class LazyCacheHolder:
        self.approx_output_change_rates = []
        self.total_steps_skipped = 0
        self.state_metadata = None
-        self.output_channels = output_channels

    def has_cache_diff(self) -> bool:
        return self.cache_diff is not None
@@ -458,7 +456,7 @@ class LazyCacheHolder:
        return self

    def clone(self):
-        return LazyCacheHolder(self.reuse_threshold, self.start_percent, self.end_percent, self.subsample_factor, self.offload_cache_diff, self.verbose, output_channels=self.output_channels)
+        return LazyCacheHolder(self.reuse_threshold, self.start_percent, self.end_percent, self.subsample_factor, self.offload_cache_diff, self.verbose)

 class LazyCacheNode(io.ComfyNode):
    @classmethod
@@ -484,7 +482,7 @@ class LazyCacheNode(io.ComfyNode):
    @classmethod
    def execute(cls, model: io.Model.Type, reuse_threshold: float, start_percent: float, end_percent: float, verbose: bool) -> io.NodeOutput:
        model = model.clone()
-        model.model_options["transformer_options"]["easycache"] = LazyCacheHolder(reuse_threshold, start_percent, end_percent, subsample_factor=8, offload_cache_diff=False, verbose=verbose, output_channels=model.model.latent_format.latent_channels)
+        model.model_options["transformer_options"]["easycache"] = LazyCacheHolder(reuse_threshold, start_percent, end_percent, subsample_factor=8, offload_cache_diff=False, verbose=verbose)
        model.add_wrapper_with_key(comfy.patcher_extension.WrappersMP.OUTER_SAMPLE, "lazycache", easycache_sample_wrapper)
        model.add_wrapper_with_key(comfy.patcher_extension.WrappersMP.PREDICT_NOISE, "lazycache", lazycache_predict_noise_wrapper)
        return io.NodeOutput(model)
--- a/comfy_extras/nodes_flipflop.py
+++ b/comfy_extras/nodes_flipflop.py
@@ -0,0 +1,43 @@
+from __future__ import annotations
+from typing_extensions import override
+
+from comfy_api.latest import ComfyExtension, io
+
+
+class FlipFlop(io.ComfyNode):
+    @classmethod
+    def define_schema(cls) -> io.Schema:
+        return io.Schema(
+            node_id="FlipFlopNew",
+            display_name="FlipFlop (New)",
+            category="_for_testing",
+            inputs=[
+                io.Model.Input(id="model"),
+                io.Float.Input(id="block_percentage", default=1.0, min=0.0, max=1.0, step=0.01),
+            ],
+            outputs=[
+                io.Model.Output()
+            ],
+            description="Apply FlipFlop transformation to model using setup_flipflop_holders method"
+        )
+
+    @classmethod
+    def execute(cls, model: io.Model.Type, block_percentage: float) -> io.NodeOutput:
+        # NOTE: this is just a hacky prototype still, this would not be exposed as a node.
+        # At the moment, this modifies the underlying model with no way to 'unpatch' it.
+        model = model.clone()
+        if not hasattr(model.model.diffusion_model, "setup_flipflop_holders"):
+            raise ValueError("Model does not have flipflop holders; FlipFlop not supported")
+        model.model.diffusion_model.setup_flipflop_holders(block_percentage)
+        return io.NodeOutput(model)
+
+class FlipFlopExtension(ComfyExtension):
+    @override
+    async def get_node_list(self) -> list[type[io.ComfyNode]]:
+        return [
+            FlipFlop,
+        ]
+
+
+async def comfy_entrypoint() -> FlipFlopExtension:
+    return FlipFlopExtension()
--- a/Show More
+++ b/Show More
Author	SHA1	Message	Date
Jedrzej Kosinski	386e854aab	Merge branch 'master' into flipflop-stream	2025-10-28 15:08:27 -07:00
Jedrzej Kosinski	61133af772	Add '--flipflop-offload' startup argument	2025-10-13 21:10:44 -07:00
Jedrzej Kosinski	586a8de8da	Merge branch 'master' into flipflop-stream	2025-10-13 21:04:37 -07:00
Jedrzej Kosinski	5329180fce	Made flipflop consider partial_unload, partial_offload, and add flip+flop to mem counters	2025-10-03 16:21:01 -07:00
Jedrzej Kosinski	0fdd327c2f	Merge branch 'master' into flipflop-stream	2025-10-03 14:32:56 -07:00
Jedrzej Kosinski	ee01002e63	Add flipflop support to (base) WAN, fix issue with applying loras to flipflop weights being done on CPU instead of GPU, left some timing functions as the lora application time could use some reduction	2025-10-02 22:02:50 -07:00
Jedrzej Kosinski	831c3cf05e	Add a temporary workaround for odd amount of blocks not producing expected results	2025-10-02 20:29:11 -07:00
Jedrzej Kosinski	0d8e8abd90	Default ro smaller blocks getting flipflopped first	2025-10-02 18:00:21 -07:00
Jedrzej Kosinski	d5001ed90e	Make flux support flipflop	2025-10-02 17:53:22 -07:00
Jedrzej Kosinski	8d7b22b720	Fixed FlipFlipModule.execute_blocks having hardcoded strings from Qwen	2025-10-02 17:49:43 -07:00
Jedrzej Kosinski	6d3ec9fcf3	Simplified flipflop setup by adding FlipFlopModule.execute_blocks helper	2025-10-02 16:46:37 -07:00
Jedrzej Kosinski	c4420b6a41	Change log string slightly	2025-10-02 15:34:35 -07:00
Jedrzej Kosinski	a282586995	Merge branch 'master' into flipflop-stream	2025-10-02 15:03:26 -07:00
Jedrzej Kosinski	0df61b5032	Fix improper index slicing for flipflop get blocks, add extra log message	2025-10-01 21:21:36 -07:00
Jedrzej Kosinski	7c896c5567	Initial automatic support for flipflop within ModelPatcher - only Qwen Image diffusion_model uses FlipFlopModule currently	2025-10-01 20:13:50 -07:00
Jedrzej Kosinski	ec156e72eb	Merge branch 'master' into flipflop-stream	2025-09-30 23:08:37 -07:00
Jedrzej Kosinski	01f4512bf8	In-progress commit on making flipflop async weight streaming native, made loaded partially/loaded completely log messages have labels because having to memorize their meaning for dev work is annoying	2025-09-30 23:08:08 -07:00
Jedrzej Kosinski	d0bd221495	Merge branch 'master' into flipflop-stream	2025-09-29 22:49:38 -07:00
Jedrzej Kosinski	8a8162e8da	Fix percentage logic, begin adding elements to ModelPatcher to track flip flop compatibility	2025-09-29 22:49:12 -07:00
Jedrzej Kosinski	ff789c8beb	Merge branch 'master' into flipflop-stream	2025-09-29 16:09:51 -07:00
Jedrzej Kosinski	0e966dcf85	Merge branch 'master' into flipflop-stream	2025-09-27 21:13:26 -07:00
Jedrzej Kosinski	6b240b0bce	Refactored old flip flop into a new implementation that allows for controlling the percentage of blocks getting flip flopped, converted nodes to v3 schema	2025-09-25 22:41:41 -07:00
Jedrzej Kosinski	f9fbf902d5	Added missing Qwen block params, further subdivided blocks function	2025-09-25 17:49:39 -07:00
Jedrzej Kosinski	f083720eb4	Refactored FlipFlopTransformer.__call__ to fully separate out actions between flip and flop	2025-09-25 16:16:51 -07:00
Jedrzej Kosinski	84e73f2aa5	Brought over flip flop prototype from contentis' fork, limiting it to only Qwen to ease the process of adapting it to be a native feature	2025-09-25 16:15:46 -07:00