Skip to content
View DevMewada1299's full-sized avatar

Highlights

  • Pro

Block or report DevMewada1299

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. GPGPU-Sim-Sensitivity-Study GPGPU-Sim-Sensitivity-Study Public

    This project evaluates sensitivity of L1 data cache size, L2 cache size, memory bandwidth (via DRAM bus width), and warp scheduler on six GPGPU-Sim benchmarks.

    Cuda

  2. gpgpu-sim_distribution-microarchitecture-changes gpgpu-sim_distribution-microarchitecture-changes Public

    Forked from gpgpu-sim/gpgpu-sim_distribution

    The dynamic SWL controller models each warp cap as an arm in a non-contextual multi-armed bandit, uses IPC-per-window as reward, first explores all caps once, and then performs a greedy hill-climb …

    C++

  3. lantern lantern Public

    A PyTorch-like deep learning tensor library in C++ for Apple Silicon from Scratch.

    C++

  4. Online-Softmax-Optimization Online-Softmax-Optimization Public

    CUDA implementation of online softmax and fused softmax+TopK for beam search, upgraded with warp shuffles, vectorized loads, mixed precision, and a bitonic Top-K network, plus benchmarks and correc…

    Cuda

  5. MCP-DOC-AUTHORING-ASSISTANT MCP-DOC-AUTHORING-ASSISTANT Public

    Python

  6. 3D-Multi-Object-Tracking-NuScenes 3D-Multi-Object-Tracking-NuScenes Public

    Python