Neev Parikh

Neev Parikh

San Francisco Bay Area
650 followers 500+ connections

About

I have a strong passion for science, math, technology, innovation, and knowledge. I have…

Activity

650 followers

See all activities

Experience

  • METR

    Berkeley, California, United States

  • -

    San Francisco Bay Area

  • -

    San Francisco Bay Area

  • -

    Providence, Rhode Island, United States

  • -

    Boston, Massachusetts, United States

  • -

    Providence County, Rhode Island, United States

  • -

    Providence, Rhode Island, United States

  • -

    Providence, Rhode Island Area

  • -

    Greater Bengaluru Area

  • -

    Providence County, Rhode Island, United States

  • -

    Bangalore

  • -

    Bangalore, India

  • -

    Bengaluru, Karnataka, India

Education

  • Brown University

    4.0

    -

    Concurrent Masters Program in CS at Brown University. Graduated magna cum laude.

  • -

  • -

    -

Licenses & Certifications

Publications

  • Learning Markov State Abstractions for Deep Reinforcement Learning

    NeurIPS 2021

    A fundamental assumption of reinforcement learning in Markov decision processes (MDPs) is that the relevant decision process is, in fact, Markov. However, when MDPs have rich observations, agents typically learn by way of an abstract state representation, and such representations are not guaranteed to preserve the Markov property. We introduce a novel set of conditions and prove that they are sufficient for learning a Markov abstract state representation. We then describe a practical training…

    A fundamental assumption of reinforcement learning in Markov decision processes (MDPs) is that the relevant decision process is, in fact, Markov. However, when MDPs have rich observations, agents typically learn by way of an abstract state representation, and such representations are not guaranteed to preserve the Markov property. We introduce a novel set of conditions and prove that they are sufficient for learning a Markov abstract state representation. We then describe a practical training procedure that combines inverse model estimation and temporal contrastive learning to learn an abstraction that approximately satisfies these conditions. Our novel training objective is compatible with both online and offline training: it does not require a reward signal, but agents can capitalize on reward information when available. We empirically evaluate our approach on a visual gridworld domain and a set of continuous control benchmarks. Our approach learns representations that capture the underlying structure of the domain and lead to improved sample efficiency over state-of-the-art deep reinforcement learning with visual features—often matching or exceeding the performance achieved with hand-designed compact state information.

    More information: https://neevparikh.com/publication/markov-abs-conference/
    PDF: https://drive.google.com/file/d/1HM_2ldTdr_5wgdfhECuj8U5zzC1YT_7g/view?usp=sharing

    See publication
  • Deep Radial-Basis Value Functions for Continuous Control

    AAAI

    A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value functions learned using a deep network with a radial-basis function (RBF) output layer. We show that the maximum action-value with respect to a deep RBVF can be easily approximated up to any desired…

    A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value functions learned using a deep network with a radial-basis function (RBF) output layer. We show that the maximum action-value with respect to a deep RBVF can be easily approximated up to any desired accuracy. Moreover, deep RBVFs can represent any true value function owing to their support for universal function approximation. We show that deep RBVFs facilitate the use of value-function-only algorithms in continuous control, and can serve as the critic in actor-critic algorithms. We extend the standard DQN algorithm to continuous control by endowing the agent with a deep RBVF, and show that it significantly outperforms value-function-only baselines and is competitive with state-of-the-art actor-critic algorithms. Together, these results reinvigorate radial-basis deep RL.

    More information: https://neevparikh.com/publication/rbfdqn-conference/
    PDF: https://ojs.aaai.org/index.php/AAAI/article/view/16828/16635

    Other authors
    See publication
  • Graph Embedding Priors for Multi-task Deep Reinforcement Learning

    NeurIPS 2020 KR2ML Workshop

    Humans appear to effortlessly generalize knowledge of similar objects and relations when learning new tasks. For example, humans playing Minecraft can learn how to use a tool to mine one block, then rapidly generalize that skill to mine others. We leverage graph-encoded object priors to capture this property and improve the performance of reinforcement learning agents across multiple tasks. We introduce a novel, flexible architecture that utilizes graph convolutional networks (GCNs), which…

    Humans appear to effortlessly generalize knowledge of similar objects and relations when learning new tasks. For example, humans playing Minecraft can learn how to use a tool to mine one block, then rapidly generalize that skill to mine others. We leverage graph-encoded object priors to capture this property and improve the performance of reinforcement learning agents across multiple tasks. We introduce a novel, flexible architecture that utilizes graph convolutional networks (GCNs), which provide a natural method to combine relational information over connected nodes. We evaluate our approach on a procedurally-generated, multi-task environment: Symbolic Procgen. Our experiments demonstrate that the method generalizes across many tasks and scales to domains with hundreds of objects and relations. Additionally, we perform ablation studies that demonstrate robustness to noisy graph priors, suggesting that the method is suitable for leveraging graphs generated from large, unstructured sources of knowledge in real-world settings.

    More information: https://neevparikh.com/publication/graph-priors-workshop/
    PDF: https://kr2ml.github.io/2020/papers/KR2ML_20_paper.pdf

    See publication
  • Locally Observable Markov Decision Process

    ICRA 2020 PAL Workshop

    Real-world robot task planning is computationally intractable in part due to the complexity of dealing with partial observability. One approach to reducing planning complexity is to assume additional model structure such as mixed-observability, factored state representations, or temporally-extended actions. We introduce a novel structured formulation, the Locally Observable Markov Decision Process, which assumes that partial observability stems from limited sensor range—objects outside sensor…

    Real-world robot task planning is computationally intractable in part due to the complexity of dealing with partial observability. One approach to reducing planning complexity is to assume additional model structure such as mixed-observability, factored state representations, or temporally-extended actions. We introduce a novel structured formulation, the Locally Observable Markov Decision Process, which assumes that partial observability stems from limited sensor range—objects outside sensor range are unobserved, but become fully observed once they are within sensor range. Plans solving tasks of this type have a specific structure: they must necessarily go through localities where objects transition from unobserved to fully observed. We introduce a novel planner that reduces planning time via a hierarchy that structures the plan around these localities, and interleaves online and offline planning. We present preliminary results in a challenging domain that shows that the locality assumption enables robots to plan effectively in the presence of this type of uncertainty.

    More information: https://neevparikh.com/publication/lomdp-workshop/
    PDF: http://cs.brown.edu/people/gdk/pubs/lomdp_ws.pdf

    See publication

Courses

  • Accelerated Intro CS

    CSCI0190

  • Analytical Mechanics

    PHYS0070

  • Computer Systems

    CSCI0330

  • Computer Vision

    CSCI1430

  • Foundations of Living Systems

    BIOL0200

  • Honors Calculus (Multivariate)

    MATH0350

  • Introduction to Music Theory

    MUSC0400A

  • Linear Algebra

    MATH0520

  • Numerical Optimization

    APMA1160

  • Principles Of Economics

    ECON0110

  • Probability and Statistics (Computing and Data Analysis)

    CSCI1450

  • TA for CSCI2951F (graduate course)

    CSCI2951F

  • The Meaning of Life

    PHIL0450

Test Scores

  • Subject SAT Math 2

    Score: 800

    800/800 in SAT Subject

  • Subject SAT Chemistry

    Score: 800

    800/800 in Subject SAT Chemistry

  • Subject SAT Physics

    Score: 800

    800/800 in Subject SAT Physics

  • SAT

    Score: 1570

    1570/1600 in the SAT

Recommendations received

View Neev’s full profile

  • See who you know in common
  • Get introduced
  • Contact Neev directly
Join to view full profile

Other similar profiles

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content

Add new skills with these courses